datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen-glm-kimi-distillation-clean
🧠 Qwen-GLM-Kimi Distillation Clean
A rigorously cleaned, finetuning-ready multi-teacher SFT corpus distilled from Qwen3.8-Max, GLM-5.2 and Kimi K3 — deduped, length-filtered and normalized for SFT with assistant-only loss.
Priorities: Quality > Cleanliness > Signal
📊 Dataset Overview
Property
Value
Total Records
57,064
Train Split
51,417 (90.1%)
Validation Split
2,833 (5.0%)
Test Split
2,814 (4.9%)
Teachers
3 (Qwen3.8-Max 47,595 /… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen-glm-kimi-distillation-clean.kimi-k3-distillation-clean
🧠 Kimi K3 Distillation — Clean
A rigorously cleaned Kimi K3-only SFT corpus of 3,653 traces — removed all 694 structurally-broken rows, normalized message schemas, merged reasoning into <think> format.
Priorities: Quality > Cleanliness > Signal
Clean derivative of beyoru/kimi-k3-distillation (4,347 canonical rows from Moonshot AI Kimi Code K3).
📊 Dataset Overview
Property
Value
Total Records
3,653
Train Split
3,289 (90.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/kimi-k3-distillation-clean.qwen3.8-max-distillation-50k-clean
🧠 Qwen3.8-Max Distillation 50K — Clean
A rigorously cleaned single-teacher SFT corpus of 49,661 traces from qwen3.8-max-preview — fixed broken <think> blocks, removed low-quality rows, added multi-format training views.
Priorities: Quality > Cleanliness > Signal
Clean derivative of r0b0tlab/qwen3.8-max-distillation-50k (49,772 rows). Companion to saidutta69/qwen-glm-kimi-distillation-clean.
📊 Dataset Overview
Property
Value
Total Records
49… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen3.8-max-distillation-50k-clean.pyine-v1-traces
PyINE-v1 Execution Traces (TACO)
This dataset contains 937,187 Python code execution traces generated by the
PyINE framework from solutions in the
TACO dataset.
Each row is a single execution trace: one code solution executed against one test input,
capturing the full sequence of variable states at every line of execution.
Dataset structure
Splits
Traces are assigned to PyINE splits at the problem level (all traces for a given problem
share the same… See the full description on the dataset page: https://huggingface.co/datasets/plstcharles-saifh/pyine-v1-traces.agentic-vibecoding-traces
🧠 Agentic Vibecoding Traces
3.5 years of real agentic coding sessions across 4 CLI agents and 25+ teacher models — fully anonymized, segmented per-task, with complete tool-call trajectories (bash commands + outputs, file edits) and chain-of-thought reasoning.
The culmination dataset: every "vibe coding" session, extracted from local agent storage, scrubbed, and packaged for SFT.
[!IMPORTANT]
Gated access. Access requests are reviewed manually. Data is anonymized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/agentic-vibecoding-traces.equational-theory-sair-note-conditioned-rollouts
Equational-theory SAIR note-conditioned rollouts
169,606 teacher rollouts across completed R0–R4 cohorts, using the original table fields/types and layout of violetxi/harvey-note-conditioned-rollouts.
Cohort
Rollouts
R0
17,408
R1
27,158
R2
33,840
R3
39,371
R4
51,829
Each cohort has one response per eligible task, sample index 0. Cohorts revisit tasks with updated note memories, so the total counts task/round instances, not distinct mathematical questions… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/equational-theory-sair-note-conditioned-rollouts.unsolved-math-clean
🧠 Unsolved Math — Clean
8,626 curated open research problems in mathematics and CS — including 122 Millennium Prize Problems — deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view.
A reasoning frontier dataset: every problem here is actually unsolved or partially solved — ideal for honest capability probing instead of contaminated benchmarks.
Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.equational-theory-sair-notes
Equational-theory SAIR notes
103,276 current usable notes, with 60,967,993 title/body tokens, from the completed R0–R4 authoritative bank.
This uses the JSONL/Parquet layout and original field names/types of
violetxi/harvey-notes-v4.
Only a recursive bank exists for this experiment; there is no inherited-only comparison bank.
Round of current revision
Notes
R0
4,870
R1
24,666
R2
26,228
R3
26,148
R4
21,364
Files and schema… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/equational-theory-sair-notes.
