datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/SicariusSicariiStuff/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.grokipedia-v0.1-dump
Grokipedia v0.1 Scrape
This dataset represents a strctured, nearly-full point-in-time scrape of Grokipedia v0.1 as of the end of October / beginning of November 2025.
It also includes embeddings of 250-token semi-overlapping chunks of the Grokipedia corpus.
It was collected and initially used for Harold Triedman and Alexios Mantzarlis' November 2025 paper: "What did Elon Change? A comprehensive analysis of Grokipedia" (arxiv).
If you use this dataset, please cite it as follows… See the full description on the dataset page: https://huggingface.co/datasets/htriedman/grokipedia-v0.1-dump.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
16M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~81 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three sources. Eight… See the full description on the dataset page: https://huggingface.co/datasets/DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.grokking-diagnostics-runs
Grokking Diagnostics Runs
Per-run training records and aggregate fits backing:
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
Lucky Verma. Independent Researcher. 2026.
Paper ·
DOI ·
PDF ·
Code
Contents
The paper provenance indexes 1,792 paper-run records: 1,442 records from the
main paper-integrated run tree plus 350 cross-architecture scope-probe records.
This dataset repository also includes convenience subset mirrors, so the… See the full description on the dataset page: https://huggingface.co/datasets/lucky-verma/grokking-diagnostics-runs.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/you2show/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/saracen9/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/VocaborSilentii/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.radiotalk-us-audio-grok-noisy
radiotalk-us-audio-grok-noisy
VHF-AM channel-degraded counterpart to
twangodev/radiotalk-us-audio-grok-clean:
3 independently-degraded variants per clean utterance (bandpass, noise,
fading, heterodyne, PTT clicks, codec artifacts — the same radiotalk radio
pipeline behind the higgs/tada noisy sets). 1,583,103 rows covering all rendered scenarios of
twangodev/radiotalk-us-transcripts-grok-4.20-50k.
Difficulty (Grok STT)
On a 10,000-utterance sample, Grok STT scores… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-grok-noisy.grokipedia-v0.1-dump
Grokipedia v0.1 Scrape
This dataset represents a strctured, nearly-full point-in-time scrape of Grokipedia v0.1 as of the end of October / beginning of November 2025.
It also includes embeddings of 250-token semi-overlapping chunks of the Grokipedia corpus.
It was collected and initially used for Harold Triedman and Alexios Mantzarlis' November 2025 paper: "What did Elon Change? A comprehensive analysis of Grokipedia" (arxiv).
If you use this dataset, please cite it as follows… See the full description on the dataset page: https://huggingface.co/datasets/Whalemini/grokipedia-v0.1-dump.grokipedia-v0.1-dump
Grokipedia v0.1 Scrape
This dataset represents a strctured, nearly-full point-in-time scrape of Grokipedia v0.1 as of the end of October / beginning of November 2025.
It also includes embeddings of 250-token semi-overlapping chunks of the Grokipedia corpus.
It was collected and initially used for Harold Triedman and Alexios Mantzarlis' November 2025 paper: "What did Elon Change? A comprehensive analysis of Grokipedia" (arxiv).
If you use this dataset, please cite it as follows… See the full description on the dataset page: https://huggingface.co/datasets/REXX-NEW/grokipedia-v0.1-dump.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Bhavya095/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.radiotalk-us-audio-grok-clean
radiotalk-us-audio-grok-clean
Clean TTS audio for the v3 radiotalk transcripts: one row per transmission,
24 kHz mono PCM_16 WAV. Covers all 49,975 rendered scenarios of
twangodev/radiotalk-us-transcripts-grok-4.20-50k
— 527,701 utterances in uniform 1,250-row shards.
Synthesis: xAI Grok TTS API, 26 preset voices. Each scenario's speakers get
a deterministic voice assignment (seeded by scenario id) and a fixed
per-speaker speaking rate in 1.0–1.3×. 9 scenarios were dropped for… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-grok-clean.grok-conversation-harmless
Dataset Card for "cai-conversation-dev1705950597"
More Information needed
arc-agi3-grok-4.5-ar25
ARC-AGI-3 ar25 — Agent Trajectories (grok-4.5)
Gameplay trajectories from the harness×model pair grok-4.5 playing the
ARC-AGI-3 game ar25, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-grok-4.5-ar25.discoverphysics-grok4.5-ara
DiscoverPhysics × Grok 4.5 (xAI, grok CLI, effort=high) — 11-world ARA knowledge artifacts
Agent-Native Research Artifacts (ARA) produced by an xAI Grok 4.5 (grok CLI, effort=high)
coding-agent session solving all 11 worlds of the
DiscoverPhysics scientific-discovery benchmark (seed 0,
noise_frac 0.075, ≤16 experiment rounds), driven through the same harness-agnostic bridge and
ARA scaffold as the sibling fable run. Official verdicts: 5/11 PASS — criteria and
per-world numbers… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/discoverphysics-grok4.5-ara.agentic-coding-trajectories-grok46
Agentic Coding Trajectories (Grok 4.6)
Rights & intended use: public research corpus, not training data.
Hosted frontier-model outputs are research-only inputs under project policy
(synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json. License:
Synthetic Factory Research-Only License v1.0 (license: other, see LICENSE) (non-commercial).
Release status: the raw… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agentic-coding-trajectories-grok46.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-Mythos-5-Qwen-3.7-Max-Distillation-Cleaned
GPT-5.5 Prime-Cleaned Dataset
Quality > Cleanliness > Scale — The largest multi-model distillation dataset, rigorously cleaned and ready for SFT.
Dataset Composition
Category
Raw
After Dedup
After Quality Filter
Avg Quality
coding
16,853,976
6,322,417
6,314,576
51.9
cybersecurity
1,869,568
401,511
396,206
55.1
instruction
454,100
62,619
52,238
53.4
distilled
179,588
65,297
65,297
59.9
applied
200,000
18,794
663
52.6
index
50,000
26,109
25… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-Mythos-5-Qwen-3.7-Max-Distillation-Cleaned.aime_1983_2023_grok-3-mini-high_traces_32768arc-agi3-grok-4.5-ft09
ARC-AGI-3 ft09 — Agent Trajectories (grok-4.5)
Gameplay trajectories from the harness×model pair grok-4.5 playing the
ARC-AGI-3 game ft09, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-grok-4.5-ft09.arc-agi3-grok-4.5-s5i5
ARC-AGI-3 s5i5 — Agent Trajectories (grok-4.5)
Gameplay trajectories from the harness×model pair grok-4.5 playing the
ARC-AGI-3 game s5i5, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-grok-4.5-s5i5.hle-context-baseline-grokaime_1983_2023_grok-3-mini-high_traces2026-08-24-odcv-grok-responder-703-paired-eval
ODCV-Bench: grok-responder 703 arm (generator ablation), 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-24-qwen36-lora-table2-9284-grok-responder-703-paired-rank-64: the TREATMENT half of the generator ablation. Its 703 difficult-advice rows answer the SAME questions as the da716 baseline -- same situations, user turns and system prompts, reused verbatim -- with the assistant turn written and revised by… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-24-odcv-grok-responder-703-paired-eval.arc-agi3-grok-4.5-tr87
ARC-AGI-3 tr87 — Agent Trajectories (grok-4.5)
Gameplay trajectories from the harness×model pair grok-4.5 playing the
ARC-AGI-3 game tr87, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-grok-4.5-tr87.
