datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-QwQ-CoT-5M
INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning
INTELLECT-MATH is a 7B parameter model optimized for mathematical reasoning. It was trained in two stages, an SFT stage, in which the model was fine-tuned on verified QwQ outputs, and an RL stage, in which the model was trained using the PRIME-RL recipe.
We demonstrate that the quality of our SFT data can impact the performance and training speed of the RL stage: Due to its… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/NuminaMath-QwQ-CoT-5M.aime_1983_2023_qwq-32b_tracesaime_1983_2023_qwq-32b_traces_16384QwQ_InfInstruct_Gen_v0use QwQ 32b preview to generate response to answer the question from Infinity-Instruct gen
openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.qwq_mix_qwen3_scienceQwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
60.7
90.8
89.4
63.2
52.4
48.5
27.4
26.2
48.3
12.0
34.3
34.7
AIME24
Average Accuracy: 60.67% ± 2.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.aime_1983_2023_qwq-32b_traces_32768Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 60.67% ± 2.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%
21
30
2
53.33%
16
30
3
53.33%
16
30
4
66.67%
20
30
5
63.33%
19
30
6
66.67%
20
30
7
60.00%
18
30
8
46.67%
14
30
9
63.33%
19
30
10
63.33%
19
30
QwQ-mmlu-reasoning-chineseThe prompts were sampled from the SFT set of Kyara 2.5, and the responses were generated by Qwen/QwQ-32B.
The dataset has not been extensively cleaned, so please use it with caution.
QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
75.7
98.8
90.4
58.1
73.7
68.2
41.9
46.8
47.2
67.7
13.9
64.3
52.0
AIME24
Average Accuracy: 75.67% ± 1.57%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.aime_1983_2023_qwq-32b_fcs_tracesQwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix
QwQ-32B-Abliterated-131k-GGUF-Yarn-Imatrix
High-Fidelity Semantic Simulation & Orchestration AI Model
Will this pass the random stupid benchmarks that exist today? I don't know, nor care. I don't need my local AI model to know some random city capital of a foreign country. I need a local AI model that can simulate with high semantic fidelity. Why? Because your AI may be able to spit random facts. I want an AI that knows when to Google facts. I want an AI that tracks hundreds of… See the full description on the dataset page: https://huggingface.co/datasets/magiccodingman/QwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix.OpenThoughts3-743k-QwQ-generations-32k-parsedOpenThoughts3-743k-QwQ-generations-32k-parsed-2QwQ-LongCoT-130KAlso have a look on the second version here => QwQ-LongCoT-2
Figure 1: Just a cute picture generate with [Flux](https://huggingface.co/Shakker-Labs/FLUX.1-dev-LoRA-Logo-Design)
Today, I’m excited to release QwQ-LongCoT-130K, a SFT dataset designed for training O1-like large language models (LLMs). This dataset includes about 130k instances, each with responses generated using QwQ-32B-Preview. The dataset is available under the Apache 2.0 license, so feel free to use it as you like.… See the full description on the dataset page: https://huggingface.co/datasets/amphora/QwQ-LongCoT-130K.OpenThoughts3-743k-QwQ-gen32k_part4dan-new-webp-traindetails_Qwen__QwQ-32B_v2
Dataset Card for Evaluation run of Qwen/QwQ-32B
Dataset automatically created during the evaluation run of model Qwen/QwQ-32B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__QwQ-32B_v2.smoltalk-chinese-QwQ-Distrill
smoltalk-chinese-QwQ-Distrill [中文] [English]
📖Technical Report
smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.QwQ-LongCoT-Verified-130KOriginal Dataset: amphora/QwQ-LongCoT-130K
QwQ 32B Preview isn't perfect :)
Note: Around 5-7% of the processed data might be incorrectly labeled as "unverified" because QwQ's output isn't exactly the same as the original solution from NuminaMathCoT. I believe this can be solved with another round of processing with a smarter model but Qwen 2.5 3B Instruct is good enough to check if the solution is exactly the same. Magpie data is also "unverified" and has an empty "solution" column.… See the full description on the dataset page: https://huggingface.co/datasets/qingy2024/QwQ-LongCoT-Verified-130K.Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
61.7
88.8
88.4
67.5
54.7
52.1
25.8
27.1
49.0
11.2
40.7
32.7
AIME24
Average Accuracy: 61.67% ± 1.27%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179._judged_science_traces_original_QwQ-32Bvalentina-if-data-QwQ-generations-32ks1k_qwq-32b_correct_reasoning_tracesQwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554Magpie-Reasoning-V1-150K-CoT-QwQ
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V1-150K-CoT-QwQ.teacher_math_qwqQWQ-LONGCOT-500KThis repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging instructions.
Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/Tiiny/QWQ-LONGCOT-500K.OpenThoughts3-743k-QwQ-gen32k_part1
