CoolFace
Datasetpublic

datajuicer/redpajama-arxiv-refined-by-data-juicer

RedPajama -- ArXiv (refined by Data-Juicer) A refined version of ArXiv dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 85GB). Dataset Information Number of samples: 1,655,259 (Keep ~95.99% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-arxiv-refined-by-data-juicer.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
2likes27downloads

datajuicer/redpajama-arxiv-refined-by-data-juicer · main · files are served by the source, never re-hosted here