frikishaan/PrimeCorpus-1B
PrimeCorpus-1B PrimeCorpus-1B is a curated 1-billion tokens text dataset created for training small and mid-scale language models. It focuses on educational, encyclopedic, and narrative domains to provide a balanced learning signal, and is created specifically for learning and experimentation. Composition Source Tokens fineweb-edu 500m finewiki 300m Gutenberg books 150m TinyStories 50m Total 1 billion Note - Token counts are measured using… See the full description on the dataset page: https://huggingface.co/datasets/frikishaan/PrimeCorpus-1B.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face