CoolFace
Datasetpublic

iNeil77/pseudo-mini-pile

A small, aggressively cleaned and de-duped pre-training corpus for academic settings. It aims to recreate something akin to The Pile but prioritizes quality for the constrained token budget academic researchers live with. It has seven config subsets and an eighth all subset that combines them for a total of ~91B tokens (GPT2 Tokenizer estimate). These splits are as follows: c4_realnews: The RealNews domain subset of the C4 dataset containing news articles. openwebtext: The OpenWebText dataset… See the full description on the dataset page: https://huggingface.co/datasets/iNeil77/pseudo-mini-pile.

sourceHugging Faceupdated 3y agoView on Hugging Face
4likes1kdownloads

iNeil77/pseudo-mini-pile · main · files are served by the source, never re-hosted here