nkandpa2/cccc_all_domains
🔓 Dolma 🍇 Creative Commons Common Crawl 🕸️ Subset of the Common Crawl corpus containing English documents with Creative Commons licenses. Snapshot Unicode Words Documents CC-MAIN-2013-20 3,851,018,197 5,529,294 CC-MAIN-2013-48 4,544,197,252 6,997,831 CC-MAIN-2014-10 4,429,217,941 6,682,672 CC-MAIN-2014-15 4,059,132,873 5,912,779 CC-MAIN-2014-23 5,193,195,765 8,253,690 CC-MAIN-2014-35 4,254,690,945 6,551,673 CC-MAIN-2014-41 4,289,814,449 6,558,170… See the full description on the dataset page: https://huggingface.co/datasets/nkandpa2/cccc_all_domains.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
