CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01codelion /sutra-1B Sutra 1B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational patterns Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.tabulartext-generation100K<n<1M2 likes1.2k downloads7mo agoHugging Face02codelion /sutra-10B Sutra 10B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 10,193,029 educational entries totaling over 10 billion tokens. This is the largest dataset in the Sutra series, designed to demonstrate that dense, curated datasets can provide best-in-class pretraining performance for small language models. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10B.tabulartext-generation1M<n<10M12 likes515 downloads7mo agoHugging Face03codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes61 downloads7mo agoHugging Face04codelion /sutra-10M Sutra 10M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 7,252 educational entries totaling approximately 10 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10M.tabulartext-generation1K<n<10K3 likes35 downloads7mo agoHugging Face05Abhisingh-18 /Sutra-1.3B-Data Sutra-1.3B — Training Data Recipe Every dataset used to build Sutra-1.3B, a 1.32B MoE model trained from scratch — with the exact config, split, text field and token share for each, plus the code that turns them into the corpus. This repo is the recipe, not the ingredients. The tokenized corpus is 93 GB of uint16 shards derived from other people's datasets, each under its own licence. Rather than redistribute that, this gives you the specification and the script — run it and you… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/Sutra-1.3B-Data.tabulartext-generationn<1K0 likes25 downloads1mo agoHugging Face06Sutranix /dsa-db2tabularn<1K0 likes2 downloads7mo agoHugging Face07Sutranix /dsa-dbtabularn<1K0 likes1 downloads7mo agoHugging Face08Sutranix /dsa-db-tinytabularn<1K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.