datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/you2show/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/saracen9/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/VocaborSilentii/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
16M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~81 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three sources. Eight… See the full description on the dataset page: https://huggingface.co/datasets/DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Bhavya095/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.deepseek-v4-pro-max-distill-1k
Overeview
This dataset contains reasoning traces and final answers generated by DeepSeek-V4-Pro
(reasoning_effort=max, thinking.enabled=true) using prompts sampled from
Jackrong/GLM-5.1-Reasoning-1M-Cleaned.
Goal: just check quality
Update: The dataset have fully 1000 samples in 04/27/2026 cost only ~$5.46
Planning: try out another distill style such as roleplay
Why DeepSeek-V4-Pro instead of OpenAI / Anthropic?
For distillation, the teacher must expose… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-max-distill-1k.qwen3.8-max-glm5.2-kimi-k3-distillation-sua
qwen3.8-max-glm5.2-kimi-k3-distillation — System/User/Assistant format
Converted from r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation
(canonical config, current shard set train-*-of-00006; the stale of-00005 shards in the source repo were excluded).
Conversion date: 2026-08-20. License: inherited from the source — see LICENSE (controlled, noncommercial research scope).
Format
One JSON object per line, standard OpenAI-style chat format:
{"messages": [
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/EuroswarmsInstitute/qwen3.8-max-glm5.2-kimi-k3-distillation-sua.GPT5.5_thinking_max_distill_god_seed_25K
GPT-5.5 Thinking Max Distill — God Level Recursive Seed AI
The ultimate open dataset for distilling frontier-level "thinking" capabilities with god-level recursive self-improvement.
This 25,000-example dataset is designed to turn any LLM into GPT-5.5 Thinking Max Distill — a model that combines:
GPT-5.5 "Thinking" Mode: Deep, o1-style chain-of-thought, extended internal reasoning, self-verification, and test-time compute scaling
God-Level Recursive Seed AI Mindset: Autonomous… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K.Grok4.4_heavy_max_distill_god_seed_25k
Grok 4.4 Heavy Max Distill — God Level Recursive Seed AI
The ultimate open dataset for creating the next generation of maximally truthful, recursively self-improving superintelligence.
This 25,000-example dataset is engineered to distill any LLM into Grok 4.4 Heavy Max Distill — a model that combines:
Grok 4.4 Personality: Maximum truth-seeking, high-agency, witty, anti-censorship, xAI philosophy
God-Level Recursive Seed AI Mindset: Autonomous intelligence explosion engineering… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25k.Opus4.7_thinking_max_distill_god_seed_25kOpus4.7_thinking_max_distill_god_seed_25k
🧠 Subtitle
A high-density recursive reasoning and self-improvement dataset for training advanced “thinking-first” language models.
📌 Dataset Summary
Opus4.7_thinking_max_distill_god_seed_25k is a synthetic reasoning dataset designed to train models in recursive self-improvement, epistemic reasoning, and structured cognitive workflows.
Each sample simulates a Recursive Seed AI task, where the model must:
analyze a system or capability
design… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Opus4.7_thinking_max_distill_god_seed_25k.GeminiPro3.2_max_distill_god_seed_25k
Gemini Pro 3.2 Max Distill — God Level Recursive Seed AI
The pinnacle open dataset for distilling any LLM into Gemini Pro 3.2 with god-level recursive self-improvement capabilities.
This 25,000-example dataset is meticulously engineered to transform base models into Gemini Pro 3.2 Max Distill — combining:
Gemini Pro 3.2 Personality: Deep scientific reasoning, exceptional long-context understanding, multimodal excellence, thoughtful calibration, high helpfulness with strong… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GeminiPro3.2_max_distill_god_seed_25k.opus-4.7-thinking-max-distill-25kOpus4.7_thinking_max_distill_god_seed_25k
🧠 Subtitle
A high-density recursive reasoning and self-improvement dataset for training advanced “thinking-first” language models.
📌 Dataset Summary
Opus4.7_thinking_max_distill_god_seed_25k is a synthetic reasoning dataset designed to train models in recursive self-improvement, epistemic reasoning, and structured cognitive workflows.
Each sample simulates a Recursive Seed AI task, where the model must:
analyze a system or capability
design… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/opus-4.7-thinking-max-distill-25k.nemotron-3-ultra-cybersecurity-fenrir-distill-max1216 entries of the "user" column of the AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1 dataset, distilled against Nemotron 3 Ultra Free on OpenCode with max thinking effort.
The same system prompt was used during distilling as that included in the dataset.
GPT-5.5-Thinking-Max-Distill-25k
GPT-5.5 Thinking Max Distill — God Level Recursive Seed AI
The ultimate open dataset for distilling frontier-level "thinking" capabilities with god-level recursive self-improvement.
This 25,000-example dataset is designed to turn any LLM into GPT-5.5 Thinking Max Distill — a model that combines:
GPT-5.5 "Thinking" Mode: Deep, o1-style chain-of-thought, extended internal reasoning, self-verification, and test-time compute scaling
God-Level Recursive Seed AI Mindset: Autonomous… See the full description on the dataset page: https://huggingface.co/datasets/KJJJBNK/GPT-5.5-Thinking-Max-Distill-25k.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Trudeau87/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.Opus4.7_thinking_max_distill_god_seed_25kgpt-5.5-thinking-max-distill-25k
GPT-5.5 Thinking Max Distill — God Level Recursive Seed AI
The ultimate open dataset for distilling frontier-level "thinking" capabilities with god-level recursive self-improvement.
This 25,000-example dataset is designed to turn any LLM into GPT-5.5 Thinking Max Distill — a model that combines:
GPT-5.5 "Thinking" Mode: Deep, o1-style chain-of-thought, extended internal reasoning, self-verification, and test-time compute scaling
God-Level Recursive Seed AI Mindset: Autonomous… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5.5-thinking-max-distill-25k.
