BramVanroy/CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
- are also present in the FineWeb or FineWeb-2 datasets;
- have no license disagreement (all found licenses have the same type; version number might differ);
- are not "non-commercial" ("nc" in license);
- are not "cc-unknown";
- do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other, high-quality resources, parsed with a better parser).
A less strict filtered version called C5f is also available that is only based around the first criterion: FineWeb data.
Created with this script.
For more information, see C5.
Progress
In the v1 release, the following crawls are included
- CC-MAIN-2019-30
- CC-MAIN-2020-05
- CC-MAIN-2023-06
- CC-MAIN-2024-51
- CC-MAIN-2024-46
- CC-MAIN-2025-05 (included in the original but not here because FineWeb and FineWeb-2 do not include it)
- CC-MAIN-2022-05
Languages
The following languages are included. This is a limited set due to computational and storage limitations.
- Afrikaans: afr
- German: deu
- English: eng
- French: fra
- Frysian: fry
- Italian: ita
- Dutch: nld
- Spanish: spa
Quantity
Counts for the strict release. Detailed number of tokens (Llama 3.3 tokenizer) and number of documents are given in the counts.json file.
Note: when comparing the original with this version, understand that: i. this version does not include crawl 2025-05 because FineWeb(-2) did not include it; ii. both 2024-46 and 2024-51 only contain English because FineWeb-2 did not include those crawls. So calculating quantity reduction after filtering should take this information into account.
