iNeil77/pseudo-mini-pile
A small, aggressively cleaned and de-duped pre-training corpus for academic settings. It aims to recreate something akin to The Pile but prioritizes quality for the constrained token budget academic researchers live with. It has seven config subsets and an eighth all subset that combines them for a total of ~91B tokens (GPT2 Tokenizer estimate). These splits are as follows: c4_realnews: The RealNews domain subset of the C4 dataset containing news articles. openwebtext: The OpenWebText dataset… See the full description on the dataset page: https://huggingface.co/datasets/iNeil77/pseudo-mini-pile.
This repository belongs to iNeil77 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
