datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc-2021-raw
cc-2021-raw
English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.cc-2020-raw
cc-2020-raw
English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.
