CoolFace
Datasetpublic

datajuicer/the-pile-philpaper-refined-by-data-juicer

The Pile -- PhilPaper (refined by Data-Juicer) A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB). Dataset Information Number of samples: 29,117 (Keep ~88.82% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes10downloads
4 commits on main
135fcac3y ago

Update README.md

hylcool
be780bd3y ago

Upload the-pile-philpaper-refine-result-preview.jsonl

hylcool
e3d1d073y ago

Update README.md

hylcool
abbe2963y ago

initial commit

hylcool