datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/SicariusSicariiStuff/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
16M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~81 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three sources. Eight… See the full description on the dataset page: https://huggingface.co/datasets/DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.Creative-Writing-Gemini3Pro-2700x
Pulitzer Diamond Prose GEMINI Seeds
This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/you2show/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/saracen9/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/VocaborSilentii/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.gemini_public_mmr1
PRISM Public SFT Data
Overview
PRISM Public SFT Data is the public supervised fine-tuning data collection used in the PRISM project.PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Before the distribution alignment and RLVR stages, we first use large-scale public multimodal demonstrations to obtain a broad SFT initialization.
This dataset serves as the public SFT data source for the… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_public_mmr1.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Bhavya095/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.gemini_distill
PRISM Gemini Distill
Overview
PRISM Gemini Distill is our self-distilled multimodal reasoning dataset collected from Gemini 3 Flash for the PRISM project.
PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline. To mitigate this issue, PRISM introduces an intermediate Distribution Alignment / Pre-alignment stage before RLVR:
SFT → Distribution Alignment / Pre-alignment → RLVR
This dataset provides high-quality Gemini 3… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_distill.Finch-Collection-Gemini-3-Flash
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the Gemini-3-Flash teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with Gemini-3-Flash as the teacher mutation… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-Gemini-3-Flash.clawloop
Nanoclaw: Verifiable Tool-Use RL
This repository packages the artifacts for Less Harness, More Signal: Efficient In-Harness RL for Autonomous Agents. The project combines:
ClawLoop, a lightweight, white-box execution loop for multi-turn tool-use rollouts;
Asymmetric Advantage Masking (AAM), which removes only positive-advantage tokens from deterministic bad-turn spans while preserving negative learning signals; and
6,970 validated workplace tasks with prompts, environment… See the full description on the dataset page: https://huggingface.co/datasets/geminiDeveloper/clawloop.HALO-Gemini-3-Flash-AppWorld
Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal)
Dataset Overview
This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem.
Field
Value
Model
Gemini 3 Flash
Benchmark
AppWorld
Split
test-normal
Total Traces
168
Total Spans
3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High
Distill
This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format.
Dataset Structure
The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.Gemini-Mental-Health-Fine-Tuning
Gemini Mental Health Fine-Tuning Dataset
A collection of curated conversational datasets prepared for experimentation with Gemini-style supervised fine-tuning and mental health chatbot development.
The datasets contain question-and-response pairs and conversational examples focused primarily on mental health topics. They also include examples designed to teach a model to decline questions that are outside the intended mental health domain.
Dataset Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/Gemini-Mental-Health-Fine-Tuning.bird-train-gemini3-flash
Dataset Card for Think2SQL-SFT
This dataset is a distilled Supervised Fine-Tuning (SFT) dataset designed to improve the reasoning capabilities of models in Text-to-SQL tasks.
It contains high-quality reasoning traces and SQL queries generated by Gemini 3 Flash.
Paper: Think2SQL: Blueprinting Reward Density and Advantage Scaling for Effective Text-To-SQL Reasoning
Base Benchmark: BIRD-Train
Dataset Description
The dataset consists of 9,428 high-quality traces, of… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-2321/bird-train-gemini3-flash.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.moltbook-obsession-gemini-flash-lite
MoltBook Obsession Experiments — Gemini Flash Lite
Multi-agent social simulation data from the Obsession experiment series on MoltBook. Each agent is given a stable, persistent real-world preoccupation (coding, fitness, hadith/commentary, forecasting, cinema) via the HEARTBEAT-v3-obsessions.md heartbeat.
This dataset captures the Gemini Flash Lite runs of the obsession heartbeat. Note that Gemini exhibits a known degeneration mode in long-running multi-agent settings: after ~15-20… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-obsession-gemini-flash-lite.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/gemini-3.1-pro-hard-high-reasoning.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
NOESIS DORA SFT Dataset
Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline.
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).
Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v14.8-NT89
Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gemini-3.1-pro-hard-high-reasoning.taubench-gemini-traces
taubench-gemini-traces
Complete HTTP-level agentic traces from running taubench_gemini benchmark tasks through an instrumented reverse proxy.
Each trace captures full request/response pairs including system prompts, user messages, assistant responses, tool calls and results, and token usage metadata.
Stats
Total sessions: 115
Multi-turn sessions (2+ LLM calls): 115
Total records: 5744
Total LLM requests: 2872
Format
Raw JSONL traces from the instrumented… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/taubench-gemini-traces.GeminiPhiDutch
Dataset Card
This dataset consists of synthetic Dutch data, in multiple styles/augmentation methods, categorized by the "type" row, this data has been filtered using Kalamazooter/DutchDatasetCleaner_Bertje.
The main motivation for creating this dataset is the lack of high-quality Dutch datasets, and the fact that existing Dutch datasets have a much smaller amount of code included compared to their English/Multilingual counterparts.
Direct Use
The dataset could be used… See the full description on the dataset page: https://huggingface.co/datasets/Kalamazooter/GeminiPhiDutch.gpt4o-coding-eval-by-gemini1_5flash-koTranslated llama-duo/gpt4o-coding-eval-by-gemini1_5flash using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
