datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M.openthoughts3-en-ar-midtrain
openthoughts3-en-ar-midtrain
Arabic translation of the OpenThoughts3_1.2M split of smoltalk2 (config Mid): long mathematical reasoning traces with <think> blocks, in a two-message user/assistant format. Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 1,135,104 source rows are present, none dropped.
The pipeline segments each message into prose and verbatim blocks (code, LaTeX, tables, and inline non-translatables are masked and never sent to the model)… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/openthoughts3-en-ar-midtrain.OpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/ESHMO-AI/OpenThoughts3-1.2M.OpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/OpenThoughts3-1.2M.glm-5.1-openthoughts3-distill
GLM-5.1-OpenThoughts3-Distill
Distilled reasoning dataset generated by GLM-5.1 from the OpenThoughts3-1.2M prompts, covering Science, Code, and Math domains.
Dataset Summary
Split
Domain
Original Prompts
Distilled (with response)
Errors
Status
Science
Physics, Chemistry, Biology, etc.
100,000
56,974
13
✅ Complete
Code
Programming, Algorithms
500,000
9,810
63,442
✅ Complete
Math
Competition Math, Proof, Algebra
850,000
1,258
181,925
✅ Complete… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/glm-5.1-openthoughts3-distill.OpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/monster75/OpenThoughts3-1.2M.OpenThoughts3-Math-17k
OpenThoughts3-Math-17k
This dataset contains 17,000 mathematical reasoning problems from the open-thoughts/OpenThoughts3-1.2M dataset, filtered for:
Math domain
Solutions with ≤16,284 tokens
Solutions containing \boxed{} answers
Format
The dataset is formatted for VERL (Versatile Reinforcement Learning) training with the following fields:
data_source: "open-thoughts/OpenThoughts3-math"
prompt: List of messages with role and content (chat format)
Includes instruction:… See the full description on the dataset page: https://huggingface.co/datasets/YYF42/OpenThoughts3-Math-17k.GLM-5.1-OpenThoughts3-Distill
GLM-5.1-OpenThoughts3-Distill
Distilled reasoning dataset generated by GLM-5.1 from the OpenThoughts3-1.2M prompts, covering Science, Code, and Math domains.
Dataset Summary
Split
Domain
Original Prompts
Distilled (with response)
Errors
Status
Science
Physics, Chemistry, Biology, etc.
100,000
56,974
13
✅ Complete
Code
Programming, Algorithms
500,000
9,810
63,442
✅ Complete
Math
Competition Math, Proof, Algebra
850,000
1,258
181,925
✅ Complete… See the full description on the dataset page: https://huggingface.co/datasets/clzoro/GLM-5.1-OpenThoughts3-Distill.openthoughts3_math
OpenThoughts3 Math
This dataset contains the math-only, complete-solution subset used for supervised fine-tuning in LLM-Fusion experiments. It was derived from open-thoughts/OpenThoughts3-1.2M.
Dataset summary
103,760 training rows
32,193 unique math questions
Up to four solutions per question, selected deterministically with seed 20260910
All rows have domain = "math" and source = "ai2-adapt-dev/openmath-2-math"
Solutions are retained only when the assistant… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-distillation/openthoughts3_math.openthoughts3_numinamath-1.5-pro_mixtureopenthoughts3-qwen3-8b-300k
OpenThoughts3 300K — Qwen3-8B SFT Trajectories
This dataset contains 300,000 synthetic reasoning trajectories generated by
Qwen/Qwen3-8B from prompts sampled
from
open-thoughts/OpenThoughts3-1.2M.
It was prepared for the supervised fine-tuning stage of the
Lightning OPD Qwen3-4B experiment, where
Qwen/Qwen3-4B-Base is the student and Qwen3-8B is the teacher.
Dataset construction
Setting
Value
Prompt source
open-thoughts/OpenThoughts3-1.2M, train split… See the full description on the dataset page: https://huggingface.co/datasets/oldpilluwu/openthoughts3-qwen3-8b-300k.openthoughts3-dedup-index
OpenThoughts3 Dedup Index
A deduplicated index over
open-thoughts/OpenThoughts3-1.2M.
The upstream dataset contains ~18× duplicate problem statements (the same
question paired with many solver trajectories). This index keeps exactly one
canonical record per unique problem, making uniform random sampling of
distinct questions trivial.
Summary
rows_scanned: 1200000
unique_questions: 65047
unique_with_gt_answer: 45622
duplicate_ratio: 18.45
domain_total_rows:
code: 250000… See the full description on the dataset page: https://huggingface.co/datasets/hyunseoki/openthoughts3-dedup-index.openthoughts3_tutor_mixturereasoning-sft-OpenThoughts3-1.2M-450K
reasoning-sft-OpenThoughts3-1.2M-450K
Converted version of open-thoughts/OpenThoughts3-1.2M, filtered to rows with exactly one valid <think>...</think> block. 750K rows were dropped due to missing or malformed think tags.
Format
Each row has three columns:
input — list of dicts (conversation turns with role and content; human → user, last gpt turn removed)
response — gpt response string including <think> reasoning block
domain — task domain (math, code, science)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-OpenThoughts3-1.2M-450K.
