G-reen/cc-2020-raw
cc-2020-raw English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.
07
