CoolFace
Datasetpublic

BramVanroy/CommonCrawl-CreativeCommons-strict

Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.

sourceHugging Faceccupdated 1y agoView on Hugging Face
2likes478downloads
Dataset Card

Common Crawl Creative Commons Corpus Strict (C5s)

A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:

  • are also present in the FineWeb or FineWeb-2 datasets;
  • have no license disagreement (all found licenses have the same type; version number might differ);
  • are not "non-commercial" ("nc" in license);
  • are not "cc-unknown";
  • do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other, high-quality resources, parsed with a better parser).

A less strict filtered version called C5f is also available that is only based around the first criterion: FineWeb data.

Created with this script.

For more information, see C5.

Progress

In the v1 release, the following crawls are included

  • CC-MAIN-2019-30
  • CC-MAIN-2020-05
  • CC-MAIN-2023-06
  • CC-MAIN-2024-51
  • CC-MAIN-2024-46
  • CC-MAIN-2025-05 (included in the original but not here because FineWeb and FineWeb-2 do not include it)
  • CC-MAIN-2022-05

Languages

The following languages are included. This is a limited set due to computational and storage limitations.

  • Afrikaans: afr
  • German: deu
  • English: eng
  • French: fra
  • Frysian: fry
  • Italian: ita
  • Dutch: nld
  • Spanish: spa

Quantity

Counts for the strict release. Detailed number of tokens (Llama 3.3 tokenizer) and number of documents are given in the counts.json file.

Note: when comparing the original with this version, understand that: i. this version does not include crawl 2025-05 because FineWeb(-2) did not include it; ii. both 2024-46 and 2024-51 only contain English because FineWeb-2 did not include those crawls. So calculating quantity reduction after filtering should take this information into account.

LanguageNo. Docs (original)**No. Docs (C5s)**No. Tokens (original)**No. Tokens (C5s)**
afr312,262350358,873,448913,178
deu9,530,74689,34011,362,859,53484,408,955
eng92,635,3727,843,16087,537,859,9587,035,305,977
fra9,234,90044,82412,366,480,02543,143,952
fry230,9101197,430,7741092
ita10,734,59768,41811,913,669,33358,765,829
nld2,827,63618,2662,757,074,70518,957,134
spa22,226,944123,30122,515,709,432113,258,753
Total147,733,3678,187,660149,009,957,2097,354,754,870