CoolFace
Datasetpublic

BramVanroy/CommonCrawl-CreativeCommons-fine

Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05 CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.

sourceHugging Faceccupdated 1y agoView on Hugging Face
5likes996downloads
Dataset Card

Common Crawl Creative Commons Corpus Fine (C5f)

A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5.

Created with this script.

For more information, see C5.

Progress

In the v1 release, the following crawls are included

  • CC-MAIN-2019-30
  • CC-MAIN-2020-05
  • CC-MAIN-2023-06
  • CC-MAIN-2024-51
  • CC-MAIN-2024-46
  • CC-MAIN-2025-05 (included in the original but not here because FineWeb and FineWeb-2 do not include it)
  • CC-MAIN-2022-05

Languages

The following languages are included. This is a limited set due to computational and storage limitations.

  • Afrikaans: afr
  • German: deu
  • English: eng
  • French: fra
  • Frysian: fry
  • Italian: ita
  • Dutch: nld
  • Spanish: spa

Quantity

Counts for the fine release. Detailed number of tokens (Llama 3.3 tokenizer) and number of documents are given in the counts.json file.

Note: when comparing the original with this version, understand that: i. this version does not include crawl 2025-05 because FineWeb(-2) did not include it; ii. both 2024-46 and 2024-51 only contain English because FineWeb-2 did not include those crawls. So calculating quantity reduction after filtering should take this information into account.

LanguageNo. Docs (original)**No. Docs (C5f)**No. Tokens (original)**No. Tokens (C5f)**
afr312,2625,753358,873,4488,214,345
deu9,530,746224,78911,362,859,534258,945,770
eng92,635,37217,528,95487,537,859,95816,629,260,476
fra9,234,900136,34912,366,480,025176,835,571
fry230,9103,240197,430,7743,879,970
ita10,734,597301,31511,913,669,333334,812,841
nld2,827,63660,5722,757,074,70560,488,015
spa22,226,944502,89622,515,709,432496,644,788
Total147,733,36718,763,868149,009,957,20917,969,081,776