datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sutra-w2c-corpus
sutra-w2c-corpus
A weights↔code training corpus for Sutra
weight→code decompilation: generated Sutra programs whose behavior is
carried by matrices, paired with those matrices (the "weights") and the
program's substrate input→output behavior. The long-term goal is a model
that maps weights → code (recovering the program from its learned
parameters).
This dataset is generated by experiments/weight_to_code_corpus.py in the
Sutra repo (where it is pinned as the corpus/ submodule)… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/sutra-w2c-corpus.sutra-1B
Sutra 1B Pretraining Dataset
A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational patterns
Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.sutra-10B
Sutra 10B Pretraining Dataset
A high-quality pedagogical dataset designed for LLM pretraining, containing 10,193,029 educational entries totaling over 10 billion tokens. This is the largest dataset in the Sutra series, designed to demonstrate that dense, curated datasets can provide best-in-class pretraining performance for small language models.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10B.sutra-100M
Sutra 100M Pretraining Dataset
A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.sutra-improved-100M
Sutra Improved 100M
A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering.
Dataset Description
This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.sutra-10M
Sutra 10M Pretraining Dataset
A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 7,252 educational entries totaling approximately 10 million tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10M.Sutra-1.3B-Data
Sutra-1.3B — Training Data Recipe
Every dataset used to build Sutra-1.3B,
a 1.32B MoE model trained from scratch — with the exact config, split, text
field and token share for each, plus the code that turns them into the corpus.
This repo is the recipe, not the ingredients. The tokenized corpus is 93 GB
of uint16 shards derived from other people's datasets, each under its own
licence. Rather than redistribute that, this gives you the specification and
the script — run it and you… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/Sutra-1.3B-Data.Jain-Sutra-Arthdsa-db2dsa-dbdsa-db-tiny
