datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
magpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedmetamath-qwen2-math
Dataset Summary
Approximately 900k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from metamath-qa https://huggingface.co/datasets/meta-math/MetaMathQA and https://huggingface.co/datasets/AI-MO/NuminaMath-CoT with only none-synthetic dataset only. We only use the prompts from metamath-qa and get response with Qwen2-math-72-instruct and rejection-sampling, the solution is filted based on the official evaluation… See the full description on the dataset page: https://huggingface.co/datasets/yingyingzhang/metamath-qwen2-math.all-Qwen2.5-72B-Instruct-AWQQwen__Qwen2.5-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-Instruct-details.bright-passage-index-gte_qwen2-1_5bQwen__Qwen2.5-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.Qwen2.5-32B-Instruct_agent_trajectories_2k
Dataset Summary
This dataset contains agent trajectories generated by the Qwen2.5-32B-Instruct model using smolagents library as the agent framework.
For more details on the method, data format, and applications, refer to the following:
Repository: https://github.com/Nardien/agent-distillation
Paper: https://arxiv.org/abs/2505.17612
JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details
Dataset Card for Evaluation run of JungZoona/T3Q-qwen2.5-14b-v1.0-e3
Dataset automatically created during the evaluation run of model JungZoona/T3Q-qwen2.5-14b-v1.0-e3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details.TSD-KD-Qwen2.5-1.5B-Instruct-Gen
TSD-KD-Qwen2.5-1.5B-Instruct-Gen
This dataset contains student-generated examples used for Token-Selective Dual Knowledge Distillation (TSD-KD), introduced in our ICLR 2026 paper:
"Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation"
Paper: https://arxiv.org/abs/2603.13260
Github: https://github.com/kmswin1/TSD-KD
Dataset Description
This dataset contains student-generated instruction-response examples from… See the full description on the dataset page: https://huggingface.co/datasets/Minsang/TSD-KD-Qwen2.5-1.5B-Instruct-Gen.MATH-PUM-qwen2.5-1.5BDataset for Process Uncertanty Model training based on the MATH dataset.
spider-rollouts-web-search-qwen2.5-7b-gaia-32-examplesQwen__Qwen2.5-32B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-32B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-32B-Instruct-details.aft-no-cot-qwen2.5-philosophy-spec
aft-no-cot-qwen2.5-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen2.5-philosophy-spec.Qwen2.5-32B-Instruct_agent_trajectories_2k_prefix
Dataset Summary
This dataset contains 2k agent trajectories generated by the Qwen2.5-32B-Instruct model using smolagents library as the agent framework.
The trajectories are collected using the "first-thought prefix" method, where each trajectory is prefixed by the model's initial reasoning steps derived from Chain-of-Thought (CoT) prompting.
For more details on the method, data format, and applications, refer to the following:
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/agent-distillation/Qwen2.5-32B-Instruct_agent_trajectories_2k_prefix.Qwen__Qwen2-1.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-1.5B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-1.5B-Instruct-details.Qwen__Qwen2-0.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-0.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-0.5B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-0.5B-Instruct-details.Qwen__Qwen2.5-Math-1.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Math-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Math-1.5B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Math-1.5B-Instruct-details.Qwen__Qwen2.5-0.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-0.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-0.5B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-0.5B-Instruct-details.llm-complex-reasoning-train-qwen2-72b-instruct-correct
Note
Data Seed from 基于封闭世界假设的复杂逻辑推理
Generate from Qwen2-72B-Instruct with prompt
train.jsonl for 推理答案和题目答案一致, no_train.jsonl推理答案和题目答案不一致
注: 题目答案不一定正确
Qwen__Qwen2-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-7B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-7B-Instruct-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.aft-cot-qwen2.5-philosophy-spec
aft-cot-qwen2.5-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen2.5-philosophy-spec.Qwen__Qwen2.5-7B-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-7B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-details.EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details
Dataset Card for Evaluation run of EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
Dataset automatically created during the evaluation run of model EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details.Qwen2.5-32B-Instruct_cot_trajectories_2k
Dataset Summary
This dataset contains CoT trajectories generated by the Qwen2.5-32B-Instruct model.
This dataset contains both correct and incorrect trajectories.
For more details on the method, data format, and applications, refer to the following:
Repository: https://github.com/Nardien/agent-distillation
Paper: https://arxiv.org/abs/2505.17612
Qwen__Qwen2-7B-details
Dataset Card for Evaluation run of Qwen/Qwen2-7B
Dataset automatically created during the evaluation run of model Qwen/Qwen2-7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-7B-details.YOYO-AI__Qwen2.5-14B-YOYO-V4-details
Dataset Card for Evaluation run of YOYO-AI/Qwen2.5-14B-YOYO-V4
Dataset automatically created during the evaluation run of model YOYO-AI/Qwen2.5-14B-YOYO-V4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/YOYO-AI__Qwen2.5-14B-YOYO-V4-details.Qwen__Qwen2-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-72B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-72B-Instruct-details.Magpie-Tanuki-Qwen2.5-72B-Answered
Magpie-Tanuki-Qwen2.5-72B-Answered
Aratako/Magpie-Tanuki-8B-annotated-96kからinput_qualityがexcellentのものを抽出し、それに対してQwen/Qwen2.5-72B-Instructで回答の再生成を行ったデータセットです。
ライセンス
基本的にはApache 2.0に準じますが、Qwen Licenseの影響を受けるため、このデータセットを使ってモデルを学習する際はこのライセンスの制約に従ってください。
