datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smoltalk-chinese-QwQ-Distrill
smoltalk-chinese-QwQ-Distrill [中文] [English]
📖Technical Report
smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.QWQ-LONGCOT-500KThis repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging instructions.
Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/Tiiny/QWQ-LONGCOT-500K.Dataset_of_Russian_thinkingRu
RTD
Описание:Russian Thinking Dataset — это набор данных, предназначенный для обучения и тестирования моделей обработки естественного языка (NLP) на русском языке. Датасет ориентирован на задачи, связанные с генерацией текста, анализом диалогов и решением математических и логических задач.
Основная информация:
Сплит: train
Количество записей: 147.046
Цели:
Обучение моделей пониманию русского языка.
Создание диалоговых систем с естественным взаимодействием.… See the full description on the dataset page: https://huggingface.co/datasets/qwqeqw/Dataset_of_Russian_thinking.Qwen__QwQ-32B-details
Dataset Card for Evaluation run of Qwen/QwQ-32B
Dataset automatically created during the evaluation run of model Qwen/QwQ-32B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__QwQ-32B-details.Qwen__QwQ-32B-Preview-details
Dataset Card for Evaluation run of Qwen/QwQ-32B-Preview
Dataset automatically created during the evaluation run of model Qwen/QwQ-32B-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__QwQ-32B-Preview-details.QwQ_32B_Preview_language_model_generation_not_confirmedbenhaotang__phi4-qwq-sky-t1-details
Dataset Card for Evaluation run of benhaotang/phi4-qwq-sky-t1
Dataset automatically created during the evaluation run of model benhaotang/phi4-qwq-sky-t1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/benhaotang__phi4-qwq-sky-t1-details.Tengentoppa-sft-seed-elyza-reasoning-QwQFINGU-AI__QwQ-Buddy-32B-Alpha-details
Dataset Card for Evaluation run of FINGU-AI/QwQ-Buddy-32B-Alpha
Dataset automatically created during the evaluation run of model FINGU-AI/QwQ-Buddy-32B-Alpha
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FINGU-AI__QwQ-Buddy-32B-Alpha-details.bunnycore__QwQen-3B-LCoT-R1-details
Dataset Card for Evaluation run of bunnycore/QwQen-3B-LCoT-R1
Dataset automatically created during the evaluation run of model bunnycore/QwQen-3B-LCoT-R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__QwQen-3B-LCoT-R1-details.OpenBuddy__openbuddy-qwq-32b-v24.2-200k-details
Dataset Card for Evaluation run of OpenBuddy/openbuddy-qwq-32b-v24.2-200k
Dataset automatically created during the evaluation run of model OpenBuddy/openbuddy-qwq-32b-v24.2-200k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenBuddy__openbuddy-qwq-32b-v24.2-200k-details.Daemontatox__Mini_QwQ-details
Dataset Card for Evaluation run of Daemontatox/Mini_QwQ
Dataset automatically created during the evaluation run of model Daemontatox/Mini_QwQ
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__Mini_QwQ-details.PromptCoT-QwQ-Dataset
Dataset Format
Each row in the dataset contains:
prompt: The input to the reasoning model, including a problem statement with the special prompting template.
completion: The expected output for supervised fine-tuning, containing a thought process wrapped in <think>...</think>, followed by the final solution.
Example
{
"prompt": "<|im_start|>user\nLet $P$ be a point on a regular $n$-gon. A marker is placed on the vertex $P$ and a random process is used to… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-QwQ-Dataset.RoboPIN-DatasetsMagpie-QwQ-Reasoning-dpoThis DPO dataset is synthesized by this way:
instructions: Magpie on QwQ-32b-Preview
outputs: QwQ-32b-Preview and calm3-22b-chat.
math-train-qwq-rs-n256math-train-qwq-rs-n192qwq_32b_factualqa_sft_dataQwQ_32B_setting_7OpenMathInstruct-2-QwQ
OpenMathInstruct-2-QwQ
Qwen/QwQ-32B-Preview solutions of 107k augmented_math problems of nvidia/OpenMathInstruct-2.
All solutions are validated and agree with the expected_answer field of OpenMathInstruct-2.
QWQ-LONGCOT-500KThis dataset is a copy of PowerInfer/QWQ-LONGCOT-500K.
This repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging… See the full description on the dataset page: https://huggingface.co/datasets/huihui-ai/QWQ-LONGCOT-500K.Tengentoppa-QwQ-reasoning-sft-elyzaqwq-MATH500-94blocksworld-mystery-4-qwq-reasoning-parts-explorationQwQ32B_GAIR_LIMO_v2prithivMLmods__QwQ-LCoT-14B-Conversational-details
Dataset Card for Evaluation run of prithivMLmods/QwQ-LCoT-14B-Conversational
Dataset automatically created during the evaluation run of model prithivMLmods/QwQ-LCoT-14B-Conversational
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__QwQ-LCoT-14B-Conversational-details.math-train-qwq-rs-n128qingy2024__QwQ-14B-Math-v0.2-details
Dataset Card for Evaluation run of qingy2024/QwQ-14B-Math-v0.2
Dataset automatically created during the evaluation run of model qingy2024/QwQ-14B-Math-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/qingy2024__QwQ-14B-Math-v0.2-details.Pinkstack__SuperThoughts-CoT-14B-16k-o1-QwQ-details
Dataset Card for Evaluation run of Pinkstack/SuperThoughts-CoT-14B-16k-o1-QwQ
Dataset automatically created during the evaluation run of model Pinkstack/SuperThoughts-CoT-14B-16k-o1-QwQ
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pinkstack__SuperThoughts-CoT-14B-16k-o1-QwQ-details.prun-QwQ-32B-MATH_train
