CoolFace
Datasetpublic

WebOrganizer/Corpus-200B

WebOrganizer/Corpus-200B [Paper] [Website] [GitHub] This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset(). The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.

sourceHugging Faceupdated 3mo agoView on Hugging Face
12likes34kdownloads
../

WebOrganizer/Corpus-200B · main · files are served by the source, never re-hosted here