WebOrganizer/Corpus-200B
WebOrganizer/Corpus-200B [Paper] [Website] [GitHub] This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset(). The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.
1235k
1version https://git-lfs.github.com/spec/v12oid sha256:6ea4ead5d1c2d0ea4518642cbf675d8be0ddc23ac026e120403b17d9c03fc1343size 348392354 