CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PrimeIntellect /NuminaMath-QwQ-CoT-5M INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning INTELLECT-MATH is a 7B parameter model optimized for mathematical reasoning. It was trained in two stages, an SFT stage, in which the model was fine-tuned on verified QwQ outputs, and an RL stage, in which the model was trained using the PRIME-RL recipe. We demonstrate that the quality of our SFT data can impact the performance and training speed of the RL stage: Due to its… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/NuminaMath-QwQ-CoT-5M.text1M<n<10M63 likes1.5k downloads2y agoHugging Face02jonathanyin /aime_1983_2023_qwq-32b_tracestabularn<1K0 likes1.2k downloads1y agoHugging Face03jonathanyin /aime_1983_2023_qwq-32b_traces_16384tabularn<1K0 likes994 downloads1y agoHugging Face04w3en2g /QwQ_InfInstruct_Gen_v0use QwQ 32b preview to generate response to answer the question from Infinity-Instruct gen texttext-generation1M<n<10M0 likes771 downloads10mo agoHugging Face05mlfoundations-dev /openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 34.3 74.5 79.4 49.4 51.0 44.3 53.9 21.5 23.1 12.2 17.0 22.7 40.1 AIME24 Average Accuracy: 34.33% ± 1.89% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.tabular10K<n<100K0 likes558 downloads1y agoHugging Face06mlfoundations-dev /qwq_mix_qwen3_sciencetabular100K<n<1M1 likes520 downloads1y agoHugging Face07mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 60.7 90.8 89.4 63.2 52.4 48.5 27.4 26.2 48.3 12.0 34.3 34.7 AIME24 Average Accuracy: 60.67% ± 2.25% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.tabular10K<n<100K1 likes451 downloads1y agoHugging Face08jonathanyin /aime_1983_2023_qwq-32b_traces_32768tabularn<1K0 likes437 downloads1y agoHugging Face09mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 Precomputed model outputs for evaluation. Evaluation Results AIME24 Average Accuracy: 60.67% ± 2.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00% 21 30 2 53.33% 16 30 3 53.33% 16 30 4 66.67% 20 30 5 63.33% 19 30 6 66.67% 20 30 7 60.00% 18 30 8 46.67% 14 30 9 63.33% 19 30 10 63.33% 19 30 tabularn<1K0 likes420 downloads1y agoHugging Face10zake7749 /QwQ-mmlu-reasoning-chineseThe prompts were sampled from the SFT set of Kyara 2.5, and the responses were generated by Qwen/QwQ-32B. The dataset has not been extensively cleaned, so please use it with caution. text1K<n<10K2 likes348 downloads2y agoHugging Face11mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 75.7 98.8 90.4 58.1 73.7 68.2 41.9 46.8 47.2 67.7 13.9 64.3 52.0 AIME24 Average Accuracy: 75.67% ± 1.57% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.tabular10K<n<100K0 likes348 downloads1y agoHugging Face12jonathanyin /aime_1983_2023_qwq-32b_fcs_tracestabularn<1K0 likes308 downloads1y agoHugging Face13magiccodingman /QwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix QwQ-32B-Abliterated-131k-GGUF-Yarn-Imatrix High-Fidelity Semantic Simulation & Orchestration AI Model Will this pass the random stupid benchmarks that exist today? I don't know, nor care. I don't need my local AI model to know some random city capital of a foreign country. I need a local AI model that can simulate with high semantic fidelity. Why? Because your AI may be able to spit random facts. I want an AI that knows when to Google facts. I want an AI that tracks hundreds of… See the full description on the dataset page: https://huggingface.co/datasets/magiccodingman/QwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix.10M<n<100M17 likes306 downloads1y agoHugging Face14jacobmorrison /OpenThoughts3-743k-QwQ-generations-32k-parsedtext100K<n<1M0 likes302 downloads1y agoHugging Face15jacobmorrison /OpenThoughts3-743k-QwQ-generations-32k-parsed-2text100K<n<1M0 likes297 downloads1y agoHugging Face16amphora /QwQ-LongCoT-130KAlso have a look on the second version here => QwQ-LongCoT-2 Figure 1: Just a cute picture generate with [Flux](https://huggingface.co/Shakker-Labs/FLUX.1-dev-LoRA-Logo-Design) Today, I’m excited to release QwQ-LongCoT-130K, a SFT dataset designed for training O1-like large language models (LLMs). This dataset includes about 130k instances, each with responses generated using QwQ-32B-Preview. The dataset is available under the Apache 2.0 license, so feel free to use it as you like.… See the full description on the dataset page: https://huggingface.co/datasets/amphora/QwQ-LongCoT-130K.texttext-generation100K<n<1M153 likes258 downloads2y agoHugging Face17jacobmorrison /OpenThoughts3-743k-QwQ-gen32k_part4text100K<n<1M0 likes208 downloads1y agoHugging Face18AI-QWQ-AI /dan-new-webp-trainimage1M<n<10M0 likes189 downloads1y agoHugging Face19OALL /details_Qwen__QwQ-32B_v2 Dataset Card for Evaluation run of Qwen/QwQ-32B Dataset automatically created during the evaluation run of model Qwen/QwQ-32B. The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__QwQ-32B_v2.text100K<n<1M0 likes181 downloads2y agoHugging Face20ChinaunicomSoftware /smoltalk-chinese-QwQ-Distrill smoltalk-chinese-QwQ-Distrill [中文] [English] 📖Technical Report smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.tabulartext-generation100K<n<1M3 likes180 downloads2y agoHugging Face21qingy2024 /QwQ-LongCoT-Verified-130KOriginal Dataset: amphora/QwQ-LongCoT-130K QwQ 32B Preview isn't perfect :) Note: Around 5-7% of the processed data might be incorrectly labeled as "unverified" because QwQ's output isn't exactly the same as the original solution from NuminaMathCoT. I believe this can be solved with another round of processing with a smarter model but Qwen 2.5 3B Instruct is good enough to check if the solution is exactly the same. Magpie data is also "unverified" and has an empty "solution" column.… See the full description on the dataset page: https://huggingface.co/datasets/qingy2024/QwQ-LongCoT-Verified-130K.text100K<n<1M31 likes169 downloads2y agoHugging Face22mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 61.7 88.8 88.4 67.5 54.7 52.1 25.8 27.1 49.0 11.2 40.7 32.7 AIME24 Average Accuracy: 61.67% ± 1.27% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179.tabular10K<n<100K0 likes156 downloads1y agoHugging Face23reasoning-proj /_judged_science_traces_original_QwQ-32Btextn<1K0 likes151 downloads1y agoHugging Face24jacobmorrison /valentina-if-data-QwQ-generations-32ktext100K<n<1M0 likes151 downloads1y agoHugging Face25jonathanyin /s1k_qwq-32b_correct_reasoning_tracestext1K<n<10K0 likes149 downloads2y agoHugging Face26mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554tabular10K<n<100K0 likes140 downloads1y agoHugging Face27Magpie-Align /Magpie-Reasoning-V1-150K-CoT-QwQ Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V1-150K-CoT-QwQ.text100K<n<1M8 likes129 downloads2y agoHugging Face28mlfoundations-dev /teacher_math_qwqtabular10K<n<100K0 likes124 downloads1y agoHugging Face29Tiiny /QWQ-LONGCOT-500KThis repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data. The dataset is available under the Apache 2.0 license. Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging instructions. Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/Tiiny/QWQ-LONGCOT-500K.text100K<n<1M124 likes121 downloads2y agoHugging Face30jacobmorrison /OpenThoughts3-743k-QwQ-gen32k_part1text100K<n<1M0 likes121 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.