BramVanroy/CommonCrawl-CreativeCommons-fine
Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05 CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.
Common Crawl Creative Commons Corpus Fine (C5f)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5.
Created with this script.
For more information, see C5.
Progress
In the v1 release, the following crawls are included
- CC-MAIN-2019-30
- CC-MAIN-2020-05
- CC-MAIN-2023-06
- CC-MAIN-2024-51
- CC-MAIN-2024-46
- CC-MAIN-2025-05 (included in the original but not here because FineWeb and FineWeb-2 do not include it)
- CC-MAIN-2022-05
Languages
The following languages are included. This is a limited set due to computational and storage limitations.
- Afrikaans: afr
- German: deu
- English: eng
- French: fra
- Frysian: fry
- Italian: ita
- Dutch: nld
- Spanish: spa
Quantity
Counts for the fine release. Detailed number of tokens (Llama 3.3 tokenizer) and number of documents are given in the counts.json file.
Note: when comparing the original with this version, understand that: i. this version does not include crawl 2025-05 because FineWeb(-2) did not include it; ii. both 2024-46 and 2024-51 only contain English because FineWeb-2 did not include those crawls. So calculating quantity reduction after filtering should take this information into account.
