nkandpa2/cccc_all_domains
π Dolma π Creative Commons Common Crawl πΈοΈ Subset of the Common Crawl corpus containing English documents with Creative Commons licenses. Snapshot Unicode Words Documents CC-MAIN-2013-20 3,851,018,197 5,529,294 CC-MAIN-2013-48 4,544,197,252 6,997,831 CC-MAIN-2014-10 4,429,217,941 6,682,672 CC-MAIN-2014-15 4,059,132,873 5,912,779 CC-MAIN-2014-23 5,193,195,765 8,253,690 CC-MAIN-2014-35 4,254,690,945 6,551,673 CC-MAIN-2014-41 4,289,814,449 6,558,170β¦ See the full description on the dataset page: https://huggingface.co/datasets/nkandpa2/cccc_all_domains.
165
No card is published for this repository, or it could not be fetched from Hugging Face right now.
