CoolFace
Datasetpublic

datajuicer/the-pile-pubmed-central-refined-by-data-juicer

The Pile -- PubMed Central (refined by Data-Juicer) A refined version of PubMed Central dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 83G). Dataset Information Number of samples: 2,694,860 (Keep ~86.96% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-central-refined-by-data-juicer.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
2likes49downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
datajuicer/the-pile-pubmed-central-refined-by-data-juicer · CoolFace