datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.CCI3-HQ
Data Description
To address the scarcity of high-quality safety datasets in the Chinese, we open-sourced the CCI (Chinese Corpora Internet) dataset on November 29, 2023.
Building on this foundation, we continue to expand the data source, adopt stricter data cleaning methods, and complete the construction of the CCI 3.0 dataset. This dataset is composed of high-quality, reliable Internet data from trusted sources.
And then with more stricter filtering, The CCI 3.0 HQ corpus… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CCI3-HQ.dolma3_mix-150B-1025
Dolma 3 Sample: 150B Mix
Dataset Sources
Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3
Source
Type
Tokens
Documents
Common Crawl
Web pages
121B (76.9%)
84.5M
olmOCR Science PDFs
Academic documents
19.9B (12.6%)
2.25M
Stack-Edu (Rebalanced)
GitHub code
11.1B (7.06%)
14.3M
arXiv
Papers with LaTeX
1.29B (0.82%)
247K
FineMath 3+
Math web pages
4.10B (2.60%)
2.57M
Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.Nemotron-SFT-Instruction-Following-Chat-v3
Dataset Description:
The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following.
The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.DR3-EvalDR3-Eval: Towards Realistic and ReproducibleDeep Research Evaluation
✨ Overview
DR³-Eval is a realistic, reproducible, and multimodal evaluation benchmark for Deep Research Agents, focusing on multi-file report generation tasks.
Existing benchmarks face a fundamental tension between realism, controllability, and reproducibility when evaluating deep research agents. DR³-Eval addresses this through the following design:… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/DR3-Eval.Nemotron-Math-Proofs-v3-SFT
Nemotron-Math-Proofs-v3-SFT
Dataset Description:
Nemotron-Math-Proofs-v3-SFT is a long-form mathematical reasoning dataset containing proof-generation, proof-refinement, verification, and meta-verification traces. The release contains 414,890 samples representing 15,818 unique problems after quality filtering.
The source pool contains 15,879 hard proof problems selected from the AoPS subset of nvidia/Nemotron-Math-Proofs-v1. Responses are generated using… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Math-Proofs-v3-SFT.UltraData-SFT-2605-no-think-8k-32k
UltraData-SFT-2605 · no_think · 8k–32k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 8k–32k range.
This is the medium-length tier intended for standard long-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.Light-R1-SFTData
Light-R1: Surpassing R1-Distill from Scratch* with $1000 through Curriculum SFT & DPO
*from models without long COT
technical report
GitHub page
Here are the two-stage SFT data we used to train Light-R1-32B.
Simply refer to stage1-76k.json and stage2-3k.json
Model
Trained From
Release Date
AIME24
AIME25
DeepSeek-R1-Distill-Llama-70B
Llama-3.3-70B-Instruct
25.1.20
70.0
54.1
DeepSeek-R1-Distill-Qwen-32B
Qwen2.5-32B
25.1.20
72.6
54.9
LIMO (32B)
Qwen2.5-32B-Instruct
25.2.4… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-R1-SFTData.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/you2show/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/saracen9/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.PKU-SafeRLHF-30K
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
Dataset Summary
The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.Nemotron-Safety-Guard-Dataset-v3
Dataset Description:
The Nemotron-Safety-Guard-Dataset-v3 (formerly known as Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It comprises approximately 514,617 samples across 12 languages: English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Mandarin, Dutch, Italian, and Korean.
This dataset is primarily synthetically generated using the CultureGuard pipeline, which… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/VocaborSilentii/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
16M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~81 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three sources. Eight… See the full description on the dataset page: https://huggingface.co/datasets/DEX9mm/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Bhavya095/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.nemotron-post-training-v2-qwen-3.5-9b-regen
Dataset Card for Nemotron Post Training v2 Qwen 3.5 9B Regen
Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using Qwen3.5 9B model.
Parameter
Value
Max Tokens
4096
Temperature
1.0
Top-k
20
Top-p
0.95
Repetition Penalty
1.5
Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled.
Category
Value
math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-qwen-3.5-9b-regen.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.qwen3.7-max-pi-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Qwen3.7 Max Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by qwen/qwen3.7-max.
JSONL files: 47
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.
Use it… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-pi-traces.UltraData-SFT-2605-no-think-32k-200k
UltraData-SFT-2605 · no_think · 32k–200k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 32k–200k range.
This is the long-context tier intended for extended-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
this repo — fxmeng/UltraData-SFT-2605-no-think-32k-200k
32k–200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-32k-200k.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.arab-dialects-20-countries-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and quality limitations.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,183,361
Dataset Server estimate: 2,064,964
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.HYDRA-M3-V0
MMM_HYDRA: Heterogeneous Yielding Dataset for Reasoning Across - Multi-hop, Multimodal, Multicompany
Dataset Description
MMM_HYDRA is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on complex financial document analysis. The dataset contains 200 carefully curated questions with answers extracted from 99 unique corporate 10-K filings across 15 industry sectors.
Key Features
Multi-Company: 54 questions (27%) span multiple… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/HYDRA-M3-V0.qwen3-4b-perfectblend-deepspec-rollout
Qwen3-4B PerfectBlend DeepSpec Rollout
This dataset contains the complete DeepSpec-aligned Qwen3-4B
self-distillation rollout over the filtered PerfectBlend corpus. The seeded
95/5 split is published as separate train and eval splits.
Splits
Split
Conversations
Shards
Path
train
1,349,860
128
data/*.jsonl
eval
71,046
64
eval/*.jsonl
total
1,420,906
192
Data construction
Canonical filtered corpus: 1,420,906 conversations.
Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Minimax M3 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by minimax/minimax-m3.
JSONL files: 31
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.
