CoolFace
Datasetpublic

WebOrganizer/Corpus-200B

WebOrganizer/Corpus-200B [Paper] [Website] [GitHub] This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset(). The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.

sourceHugging Faceupdated 3mo agoView on Hugging Face
12likes35kdownloads
CC_shard_00000053_processed.jsonl.zst4 linesDownload Raw Back to documents
1version https://git-lfs.github.com/spec/v12oid sha256:6ea4ead5d1c2d0ea4518642cbf675d8be0ddc23ac026e120403b17d9c03fc1343size 348392354 
WebOrganizer/Corpus-200B · CoolFace