datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime_1983_2023_qwq-32b_tracesaime_1983_2023_qwq-32b_traces_16384openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.qwq_mix_qwen3_scienceQwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
60.7
90.8
89.4
63.2
52.4
48.5
27.4
26.2
48.3
12.0
34.3
34.7
AIME24
Average Accuracy: 60.67% ± 2.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.aime_1983_2023_qwq-32b_traces_32768Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 60.67% ± 2.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%
21
30
2
53.33%
16
30
3
53.33%
16
30
4
66.67%
20
30
5
63.33%
19
30
6
66.67%
20
30
7
60.00%
18
30
8
46.67%
14
30
9
63.33%
19
30
10
63.33%
19
30
QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
75.7
98.8
90.4
58.1
73.7
68.2
41.9
46.8
47.2
67.7
13.9
64.3
52.0
AIME24
Average Accuracy: 75.67% ± 1.57%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.aime_1983_2023_qwq-32b_fcs_tracessmoltalk-chinese-QwQ-Distrill
smoltalk-chinese-QwQ-Distrill [中文] [English]
📖Technical Report
smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
61.7
88.8
88.4
67.5
54.7
52.1
25.8
27.1
49.0
11.2
40.7
32.7
AIME24
Average Accuracy: 61.67% ± 1.27%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179.QwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554teacher_math_qwqQwen2.5-7B-Instruct_openthoughts3_math_100k_annotated_QwQ-32B_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_openthoughts3_math_100k_annotated_QwQ-32B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
46.3
86.2
87.4
59.5
48.5
25.1
7.9
8.5
35.0
11.0
17.9
24.7
AIME24
Average Accuracy: 46.33% ± 1.91%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_openthoughts3_math_100k_annotated_QwQ-32B_eval_8179.QWQ_bench_mmlu_pro_distilled_r1_styleteacher_code_qwqQWQ_bench_mmlu_pro_distilledQwen__QwQ-32B-details
Dataset Card for Evaluation run of Qwen/QwQ-32B
Dataset automatically created during the evaluation run of model Qwen/QwQ-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__QwQ-32B-details.Qwen__QwQ-32B-Preview-details
Dataset Card for Evaluation run of Qwen/QwQ-32B-Preview
Dataset automatically created during the evaluation run of model Qwen/QwQ-32B-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__QwQ-32B-Preview-details.phi_30K_qwq_0K_eval_2e29
mlfoundations-dev/phi_30K_qwq_0K_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
26.7
42.2
44.8
8.1
32.0
47.1
0.8
0.3
0.1
20.7
2.3
0.3
AIME24
Average Accuracy: 26.67% ± 1.63%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
26.67%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_30K_qwq_0K_eval_2e29.dpo-base-100k-qwq-judge-random-rejectedQwQ_Benchmark_Distill_sharegptoriginal qwq distilled without gpqa
benhaotang__phi4-qwq-sky-t1-details
Dataset Card for Evaluation run of benhaotang/phi4-qwq-sky-t1
Dataset automatically created during the evaluation run of model benhaotang/phi4-qwq-sky-t1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/benhaotang__phi4-qwq-sky-t1-details.FINGU-AI__QwQ-Buddy-32B-Alpha-details
Dataset Card for Evaluation run of FINGU-AI/QwQ-Buddy-32B-Alpha
Dataset automatically created during the evaluation run of model FINGU-AI/QwQ-Buddy-32B-Alpha
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FINGU-AI__QwQ-Buddy-32B-Alpha-details.bunnycore__QwQen-3B-LCoT-R1-details
Dataset Card for Evaluation run of bunnycore/QwQen-3B-LCoT-R1
Dataset automatically created during the evaluation run of model bunnycore/QwQen-3B-LCoT-R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__QwQen-3B-LCoT-R1-details.OpenBuddy__openbuddy-qwq-32b-v24.2-200k-details
Dataset Card for Evaluation run of OpenBuddy/openbuddy-qwq-32b-v24.2-200k
Dataset automatically created during the evaluation run of model OpenBuddy/openbuddy-qwq-32b-v24.2-200k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenBuddy__openbuddy-qwq-32b-v24.2-200k-details.e1_science_longest_qwq_togetherteacher_science_qwqDaemontatox__Mini_QwQ-details
Dataset Card for Evaluation run of Daemontatox/Mini_QwQ
Dataset automatically created during the evaluation run of model Daemontatox/Mini_QwQ
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__Mini_QwQ-details.QwQ-32B-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completions
