datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uno-Curriculum
Uno-Curriculum
Training corpus for a hierarchical-delegation router: a small language
model that decomposes a task into subtasks and routes each subtask to a
(worker model, skill) pair.
Every row comes from a real public HuggingFace dataset — the
question and gold_answer are sampled verbatim from the dataset
identified by the source field. Every row then goes through the
same three-stage pipeline (router probe → teacher trajectory →
noise removal) to obtain the multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/tinaxie/Uno-Curriculum.ru-en-code-curriculum
RuEn Code Curriculum
RuEn Code Curriculum is a curated Russian-English dataset for continued pretraining (CPT) and supervised fine-tuning (SFT) of small code-oriented language models.
This public release contains only records classified as redistributable. Local-training-only web and code sources used by the internal curriculum are intentionally excluded.
Dataset summary
Configuration
Split
Records
Tokens
sft
train
53,278
13,997,239
sft
reserve
19,225… See the full description on the dataset page: https://huggingface.co/datasets/sup2ch/ru-en-code-curriculum.adaptive-curriculum-tool-calling-pool
Adaptive Curriculum Tool-Calling Pool — v2.1
A gated snapshot of the tool-calling training-data pool produced by the
Adaptive Curriculum for Tool Calling sub-experiment. This is the additive v2.1
revision: it keeps the entire v1 + v2 payload and adds the six per-campaign
partition manifests under manifests/partitions/. Nothing from v1 or v2 was
re-encoded, recompressed, moved, or rewritten.
Access is gated. The repository uses manual gating. You must be granted access
by the… See the full description on the dataset page: https://huggingface.co/datasets/kesava89/adaptive-curriculum-tool-calling-pool.
