CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jonathanyin /aime_1983_2023_qwq-32b_tracestabularn<1K0 likes1.2k downloads1y agoHugging Face02jonathanyin /aime_1983_2023_qwq-32b_traces_16384tabularn<1K0 likes1.1k downloads1y agoHugging Face03mlfoundations-dev /openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 34.3 74.5 79.4 49.4 51.0 44.3 53.9 21.5 23.1 12.2 17.0 22.7 40.1 AIME24 Average Accuracy: 34.33% ± 1.89% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.tabular10K<n<100K0 likes549 downloads1y agoHugging Face04mlfoundations-dev /qwq_mix_qwen3_sciencetabular100K<n<1M1 likes527 downloads1y agoHugging Face05jonathanyin /aime_1983_2023_qwq-32b_traces_32768tabularn<1K0 likes447 downloads1y agoHugging Face06mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 Precomputed model outputs for evaluation. Evaluation Results AIME24 Average Accuracy: 60.67% ± 2.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00% 21 30 2 53.33% 16 30 3 53.33% 16 30 4 66.67% 20 30 5 63.33% 19 30 6 66.67% 20 30 7 60.00% 18 30 8 46.67% 14 30 9 63.33% 19 30 10 63.33% 19 30 tabularn<1K0 likes420 downloads1y agoHugging Face07mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 75.7 98.8 90.4 58.1 73.7 68.2 41.9 46.8 47.2 67.7 13.9 64.3 52.0 AIME24 Average Accuracy: 75.67% ± 1.57% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.tabular10K<n<100K0 likes347 downloads1y agoHugging Face08mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 60.7 90.8 89.4 63.2 52.4 48.5 27.4 26.2 48.3 12.0 34.3 34.7 AIME24 Average Accuracy: 60.67% ± 2.25% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.tabular10K<n<100K1 likes332 downloads1y agoHugging Face09jonathanyin /aime_1983_2023_qwq-32b_fcs_tracestabularn<1K0 likes329 downloads1y agoHugging Face10ChinaunicomSoftware /smoltalk-chinese-QwQ-Distrill smoltalk-chinese-QwQ-Distrill [中文] [English] 📖Technical Report smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.tabulartext-generation100K<n<1M3 likes159 downloads2y agoHugging Face11mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554tabular10K<n<100K0 likes142 downloads1y agoHugging Face12mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 61.7 88.8 88.4 67.5 54.7 52.1 25.8 27.1 49.0 11.2 40.7 32.7 AIME24 Average Accuracy: 61.67% ± 1.27% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179.tabular10K<n<100K0 likes142 downloads1y agoHugging Face13mlfoundations-dev /teacher_math_qwqtabular10K<n<100K0 likes124 downloads1y agoHugging Face14mlfoundations-dev /Qwen2.5-7B-Instruct_openthoughts3_math_100k_annotated_QwQ-32B_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_openthoughts3_math_100k_annotated_QwQ-32B_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 46.3 86.2 87.4 59.5 48.5 25.1 7.9 8.5 35.0 11.0 17.9 24.7 AIME24 Average Accuracy: 46.33% ± 1.91% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_openthoughts3_math_100k_annotated_QwQ-32B_eval_8179.tabular10K<n<100K0 likes113 downloads1y agoHugging Face15reasoningMIA /QWQ_bench_mmlu_pro_distilled_r1_styletabularn<1K0 likes106 downloads11mo agoHugging Face16mlfoundations-dev /teacher_code_qwqtabular10K<n<100K0 likes97 downloads1y agoHugging Face17reasoningMIA /QWQ_bench_mmlu_pro_distilledtabularn<1K0 likes84 downloads11mo agoHugging Face18open-llm-leaderboard /Qwen__QwQ-32B-detailsgated Dataset Card for Evaluation run of Qwen/QwQ-32B Dataset automatically created during the evaluation run of model Qwen/QwQ-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__QwQ-32B-details.tabular10K<n<100K0 likes69 downloads2y agoHugging Face19open-llm-leaderboard /Qwen__QwQ-32B-Preview-detailsgated Dataset Card for Evaluation run of Qwen/QwQ-32B-Preview Dataset automatically created during the evaluation run of model Qwen/QwQ-32B-Preview The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__QwQ-32B-Preview-details.tabular10K<n<100K0 likes56 downloads2y agoHugging Face20allenai /dpo-base-100k-qwq-judge-random-rejectedtabular100K<n<1M2 likes53 downloads1y agoHugging Face21reasoningMIA /QwQ_Benchmark_Distill_sharegptoriginal qwq distilled without gpqa tabular1K<n<10K1 likes50 downloads1y agoHugging Face22open-llm-leaderboard /benhaotang__phi4-qwq-sky-t1-detailsgated Dataset Card for Evaluation run of benhaotang/phi4-qwq-sky-t1 Dataset automatically created during the evaluation run of model benhaotang/phi4-qwq-sky-t1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/benhaotang__phi4-qwq-sky-t1-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face23open-llm-leaderboard /FINGU-AI__QwQ-Buddy-32B-Alpha-detailsgated Dataset Card for Evaluation run of FINGU-AI/QwQ-Buddy-32B-Alpha Dataset automatically created during the evaluation run of model FINGU-AI/QwQ-Buddy-32B-Alpha The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FINGU-AI__QwQ-Buddy-32B-Alpha-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face24mlfoundations-dev /phi_30K_qwq_0K_eval_2e29 mlfoundations-dev/phi_30K_qwq_0K_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 26.7 42.2 44.8 8.1 32.0 47.1 0.8 0.3 0.1 20.7 2.3 0.3 AIME24 Average Accuracy: 26.67% ± 1.63% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 23.33% 7 30 2 26.67%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_30K_qwq_0K_eval_2e29.tabular10K<n<100K0 likes46 downloads1y agoHugging Face25open-llm-leaderboard /bunnycore__QwQen-3B-LCoT-R1-detailsgated Dataset Card for Evaluation run of bunnycore/QwQen-3B-LCoT-R1 Dataset automatically created during the evaluation run of model bunnycore/QwQen-3B-LCoT-R1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__QwQen-3B-LCoT-R1-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face26dmitriihook /qwq-32b-planning-6-blocks-self-probing-state-distilabel Dataset Card for qwq-32b-planning-6-blocks-self-probing-state-distilabel This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-6-blocks-self-probing-state-distilabel.tabular10K<n<100K0 likes44 downloads2y agoHugging Face27open-llm-leaderboard /OpenBuddy__openbuddy-qwq-32b-v24.2-200k-detailsgated Dataset Card for Evaluation run of OpenBuddy/openbuddy-qwq-32b-v24.2-200k Dataset automatically created during the evaluation run of model OpenBuddy/openbuddy-qwq-32b-v24.2-200k The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenBuddy__openbuddy-qwq-32b-v24.2-200k-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face28mlfoundations-dev /e1_science_longest_qwq_togethertabular10K<n<100K0 likes43 downloads1y agoHugging Face29mlfoundations-dev /teacher_science_qwqtabular10K<n<100K0 likes40 downloads1y agoHugging Face30mothnaZl /QwQ-32B-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completionstabularn<1K0 likes38 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.