CoolFace
Datasetpublic

G-reen/cc-2020-raw

cc-2020-raw English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes7downloads

G-reen/cc-2020-raw · main · files are served by the source, never re-hosted here