datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
16M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~81 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three sources. Eight… See the full description on the dataset page: https://huggingface.co/datasets/DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/you2show/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/saracen9/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/VocaborSilentii/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Bhavya095/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.kimi-k3-open-swe-distillation
Kimi K3 Open-SWE Distillation
Sanitized action-window data from a black-box Kimi K3 distillation experiment over NVIDIA Open-SWE-Traces. The rows retain native messages, exact tool schemas in tools, and metadata identifying the supervised assistant action.
Configurations
exact_replay: 69 action windows across 35 projects where moonshotai/kimi-k3, queried through OpenRouter, reproduced the source target action exactly.
critic_accepted: 198 source action windows… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kimi-k3-open-swe-distillation.self-self-distillation
self-self-distillation
Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation,
computed on the sky_work_math subset of
PrimeIntellect/SYNTHETIC-2-RL with
Qwen/Qwen3-4B.
For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at
identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the
per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.aicivs-npc-distillation
AICivs NPC distillation data
The data behind the AICivs student: teacher answers for the nine
operations of the AICivs service contract (a Minecraft mod whose villages are living civilizations), filtered the way the
game validates them, and formatted as the exact rows the student was trained on. Everything is synthetic: requests were
sampled from a 12-civilization world simulated headless for 120 seasons (souls, memories, chronicles, quest manifests),
answers were written by… See the full description on the dataset page: https://huggingface.co/datasets/cow9000/aicivs-npc-distillation.qwen3.8-max-glm5.2-kimi-k3-distillation-sua
qwen3.8-max-glm5.2-kimi-k3-distillation — System/User/Assistant format
Converted from r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation
(canonical config, current shard set train-*-of-00006; the stale of-00005 shards in the source repo were excluded).
Conversion date: 2026-08-20. License: inherited from the source — see LICENSE (controlled, noncommercial research scope).
Format
One JSON object per line, standard OpenAI-style chat format:
{"messages": [
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/EuroswarmsInstitute/qwen3.8-max-glm5.2-kimi-k3-distillation-sua.AtmosphericQA-1k-Chinese-Distillation
AtmosphericQA-distillation
Overview
AtmosphericQA-distillation is a Chinese supervised fine-tuning (SFT) question–answering dataset focused on atmospheric science and meteorology.The dataset is constructed via knowledge distillation from the Gemini 3 Flash Preview model, with the goal of providing systematic, structured, and domain-specific scientific knowledge for Chinese large language models.
It covers a broad range of subfields, from fundamental atmospheric theory to… See the full description on the dataset page: https://huggingface.co/datasets/phoenixcph/AtmosphericQA-1k-Chinese-Distillation.qwen3.6-27b-self-data-distillation-dataset
Qwen3.6-27B Self-Data-Distillation Trajectories
Single‑turn reasoning trajectories generated by running Qwen3.6‑27B (via vLLM). Each trajectory contains a system prompt, a user task, and the model's full output (including reasoning steps embedded in the assistant content field).
Data Format
Four JSONL files, one per category. Each line is:
{
"id": "traj_<timestamp>_<idx>_<seq>",
"source": "synthetic-qwen3.6-27b",
"task": "<the prompt given to the model>"… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/qwen3.6-27b-self-data-distillation-dataset.faithful-tom-distillation
Dataset for "Faithful Theory of Mind Distillation"
This repository contains the datasets used in the paper "Faithful Theory of Mind Distillation: Why Preference Based Refinement Improves Imitation", accepted to the AAAI 2026 ToM4AI Workshop.
Dataset Structure
The repository contains two files corresponding to the two training stages described in the paper:
train_sft_combined.jsonl:
Purpose: Used for the Supervised Fine-Tuning (SFT) stage.
Content: Contains social… See the full description on the dataset page: https://huggingface.co/datasets/ArpitSinghGautam/faithful-tom-distillation.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample shows what a training-ready video understanding distillation dataset can look like.
Why this exists
Most teams evaluating outside data vendors want to know one thing first:
What does the delivered data actually look like?
This sample is designed to answer that question.
It demonstrates how raw video clips can be converted into structured, model-ready supervision for:
video understanding
multimodal SFT… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/video-understanding-distillation-sample.video-understanding-distillation-sample
Video Understanding Distillation Sample
This public sample demonstrates what a training-ready video understanding / multimodal distillation dataset can look like.
Intended purpose
This dataset is not a production corpus. It is a schema demonstration for potential partners evaluating SuperviseLab's delivery approach.
What it shows
clip-level metadata
short and long captions
OCR text
transcript
speaker attribution
structured JSON targets
distillation-ready… See the full description on the dataset page: https://huggingface.co/datasets/metavi/video-understanding-distillation-sample.finance-slm-distillation-data
Finance Instruction SFT Dataset
Overview
This repository provides a curated finance-oriented instruction dataset designed for conversational supervised fine-tuning (SFT) of large language models.
The dataset is derived from:
heladell/Finance_DeepSeek-R1-Distill-dataset
and was further processed through a custom preprocessing pipeline focused on:
finance-domain filtering
reasoning-quality filtering
answer-quality filtering
conversational formatting
train/test split… See the full description on the dataset page: https://huggingface.co/datasets/ash0t/finance-slm-distillation-data.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Trudeau87/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.distillation01
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
distillation01
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
text-to-sql-struct-distillation-minidev
结构化 Text-to-SQL 蒸馏 Mini-Dev 派生 SFT
本仓库发布由 BIRD Mini-Dev 500 个样本构造的派生 SFT messages 数据,共 500 条。
文件与边界
minidev_sft_messages.jsonl:500 条 messages 格式的派生 SFT 样本。
不含 Mini-Dev SQLite 数据库、gold SQL、原始 schema 文件或上游数据库内容。
评测中使用 BIRD 作者提供的 SQLite 集合型 Execution Accuracy (EX) 定义;该评测代码与复现说明在 GitHub 工程中维护。
来源、署名与许可证
本数据为 BIRD Mini-Dev 的派生文本内容。上游仓库:https://github.com/bird-bench/mini_dev。上游 README 标明 CC BY-SA 4.0,因此本仓库按 CC BY-SA 4.0 发布。使用或再发布时请保留对 BIRD 与 Mini-Dev… See the full description on the dataset page: https://huggingface.co/datasets/craboy4/text-to-sql-struct-distillation-minidev.
