datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M.OpenThoughts3openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.OpenThoughts3-full-filtered-mathopenthoughts3-en-ar-midtrain
openthoughts3-en-ar-midtrain
Arabic translation of the OpenThoughts3_1.2M split of smoltalk2 (config Mid): long mathematical reasoning traces with <think> blocks, in a two-message user/assistant format. Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 1,135,104 source rows are present, none dropped.
The pipeline segments each message into prose and verbatim blocks (code, LaTeX, tables, and inline non-translatables are masked and never sent to the model)… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/openthoughts3-en-ar-midtrain.openthoughts3-unfinished-promptsOpenThoughts3-456k-no-cotOpenThoughts3-1.2M
paper |
dataset |
model
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
OpenThoughts3-1.2M
Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀
OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M.
This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/ESHMO-AI/OpenThoughts3-1.2M.QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
75.7
98.8
90.4
58.1
73.7
68.2
41.9
46.8
47.2
67.7
13.9
64.3
52.0
AIME24
Average Accuracy: 75.67% ± 1.57%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.OpenThoughts3-456k-no-cot-with-olmo-system-promptOpenThoughts3-1.2M-no-cotOpenThoughts3-456k-gpt4.1-cotOpenThoughts3-1.2MOpenThoughts3-743k-QwQ-generations-32k-parsedOpenThoughts3-full-filtered-codeOpenThoughts3-743k-QwQ-generations-32k-parsed-2OpenThoughts3-456kopenthoughts3_100k_llama3_eval_5554
mlfoundations-dev/openthoughts3_100k_llama3_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
37.0
75.2
83.8
11.2
45.2
45.1
44.4
13.8
18.3
9.7
19.3
30.3
31.9
AIME24
Average Accuracy: 37.00% ± 1.29%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_100k_llama3_eval_5554.OpenThoughts3-743k-QwQ-gen32k_part4OpenThoughts3-mathOpenThoughts3-math-boxed-326kopenthoughts3_300k_annotated_Qwen3-32Bopenthoughts-30kopenthoughts3_herorun_ckpt06500_eval_27e9
mlfoundations-dev/openthoughts3_herorun_ckpt06500_eval_27e9
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 59.95% ± 0.79%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
59.30%
303
511
2
62.82%
321
511
3
60.08%
307
511
4
58.51%
299
511
5
61.45%
314
511
6
57.53%
294
511
OpenThoughts3-1.2M-200k-sftQwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554openthoughts3_100k_qwen25_1b_bsz256_lr4e5_epochs5_eval_8179
mlfoundations-dev/openthoughts3_100k_qwen25_1b_bsz256_lr4e5_epochs5_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
6.0
37.5
57.4
16.1
25.4
6.5
2.0
3.1
9.7
10.3
3.6
2.7
AIME24
Average Accuracy: 6.00% ± 0.79%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_100k_qwen25_1b_bsz256_lr4e5_epochs5_eval_8179.DeepSeek-R1-Distill-Qwen-7B_OpenThoughts3_eval_8179
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_OpenThoughts3_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
64.0
91.2
89.0
65.4
47.3
61.8
23.3
24.2
51.3
10.9
45.5
35.3
AIME24
Average Accuracy: 64.00% ± 1.23%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_OpenThoughts3_eval_8179.OpenThoughts3-full-filtered-science-no-cot-decontam-v2OpenThoughts3-743k-QwQ-gen32k_part1
