datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-2M-Thinking.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0).
🎯 Key Highlights
123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.CodeX-7M-Non-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is curated from high-quality public sources and enhanced with synthetic data from both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-7M-Non-Thinking.thinking-benchmark-90
Thinking Benchmark
A calibration pool of 90 competition-mathematics problems assembled to study how output / reasoning-trace length varies with problem difficulty across frontier language models. Part of the Cost of Overthinking research project.
Dataset at a glance
Source
n
Difficulty
Contamination risk
AIME 2026
29
3–5
low
OlymMATH
41
4–6
medium
HMMT February 2026
12
4–5
low
MATH-500
5
2–3
high
FrontierMath-style
3
6
medium
Difficulty is… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-90.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/adrianmele/CodeX-2M-Thinking.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/dans25275/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.arxiv-qa-thinking
ArXiv Q&A with Thinking Dataset
This dataset contains question-answer pairs generated by MiniMax-M2.1 based on academic articles from PursuitOfDataScience/arxiv-llama4-maverick-abstract.
Dataset Description
For each academic article, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Question: An insightful question testing understanding of key concepts
Answer: A detailed answer based on the article content
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/arxiv-qa-thinking.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/CodeX-2M-Thinking.sciqa-thinking
sciqa-thinking
Randomly extracted 3000 rows from sciq and prompting Qwen3-14b to generate the intermediate reasoning traces, we created this dataset.
This should be used for LLM post-training, especially RL.
CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/CodeX-2M-Thinking.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.Adaptive_Skip_thinking_Reasoningannotations_creators:
found
language_creators:
llm-generated
languages:
ru
licenses:
unknown
multilinguality:
monolingual
pretty_name: Cerebras Adaptive Reasoning (Russian)
size_categories:
n-examples--1K
source_datasets: []
task_categories:
text-generation
reasoning
task_ids:
chain-of-thought
program-of-thought
skip-thinking
papers_with_code:
null
train_eval_split: []
configs:
default
Описание Датасета
cerebras-adaptive-reasoning-ru — это синтетический датасет, предназначенный для… See the full description on the dataset page: https://huggingface.co/datasets/Siesher/Adaptive_Skip_thinking_Reasoning.gsm8k-thinking
GSM8K Thinking
This dataset contains responses generated by MiniMax-M2.1 for math word problems from the openai/gsm8k dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Train Examples
7,473
Test Examples
1,319
Total Examples
8,792
Total Tokens
10,506,774
Avg Tokens/Example
1,195
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/gsm8k-thinking.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096
Derived dataset note
This dataset was derived from OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking as a part of arxiv.org/abs/2603.22276.
Field changes:
question -> query
qwen3vl_235b_thinking_response -> response
image -> images (single-item list)
added tok_len, computed with tokenizer Qwen/Qwen3-8B on query + '\n\n' + response
add_special_tokens=False
The original README content is preserved below.
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/eyes-ml/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096.ThinkingData-200K-Turkish
Dataset Card for ThinkingData-200K-Turkish
Language: Turkish
Dataset Description
This repository contains a dataset for Turkish version of the Deepseek 1.5B model. The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈205K
Translation tool: Google Translate
Data format: Prompt, Think, Response
Kimi-K2.6-Thinking-200x
Dataset Card (Kimi-K2.6-Thinking-200x)
Dataset Summary
Kimi-K2.6-Reasoning-207 is a high-quality distilled reasoning dataset designed for supervised fine-tuning (SFT) of small language models.
This dataset uses a curated seed question set covering Mathematics, Code, Logic, Science, Analysis, and Instruction-following domains. By calling the Kimi-K2.6 model via the Moonshot AI API as the teacher model, it generates high-quality responses featuring long-form step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/uniquealexx/Kimi-K2.6-Thinking-200x.agentmujo-thinking
agentmujo-thinking (v0.1.0)
Uzorci sa eksplicitnim <think> tragovima razmišljanja — obrazac:
opservacija → procjena rizika → odluka → akcija/odbijanje. Namjena:
thinking-alignment nastavak treninga (model u thinking režimu mora
razmišljati PRIJE djelovanja, posebno kod sigurnosnih odluka).
Uzoraka: 25 (15 function-calling + 10 agentic-terminal)
Format: JSONL, ista schema; <think> blokovi su doslovni tekst
unutar assistant poruka (nativni Qwen format).
Kvalitet: 25/25 ACCEPT… See the full description on the dataset page: https://huggingface.co/datasets/shaban2024/agentmujo-thinking.ThinkingData-425K-Turkish
Dataset Card for Instruction-425K-Turkish
Language: Turkish
Dataset Description
The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈425K
Translation tool: Google Translate
Data format: Prompt, Reasoning, Response
CodeX-Thinking-Gemma-4-31B-ITAll prompts were taken from Modotte/CodeX-2M-Thinking, which contains multiple traces per prompt whereas this dataset only provides one trace per prompt. Generations were with https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4 (a mix of BF16/FP8 weights that NVIDIA configured with FP8 KV cache; benchmarks show performs similarly to BF16 for coding). No system prompt was used.
thinking-benchmark-hard-but-doable-10
Thinking Benchmark — Hard-but-Doable (10-problem panel)
10 problems solved 8/8 at k=8 across the model panel (non-trivial trace length, but reliably solvable). Selected for measuring trace length variance on problems near ceiling, as a companion to the true edge-of-capability sets.
Subset of tyrtleli/thinking-benchmark-90.
Note: the problem (full text) column is not included here — join by id against tyrtleli/thinking-benchmark-90 before running k=32.
Problems… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-hard-but-doable-10.toucan-agentic-thinking
Toucan Agentic with Thinking Dataset
This dataset contains agentic reasoning responses generated by MiniMax-M2.1 based on questions from Agent-Ark/Toucan-1.5M_SFT.
Dataset Description
For each user question, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Response: A complete, helpful answer in natural language
The original tool definitions are preserved in the tools field for reference.
Statistics
Split
Examples… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/toucan-agentic-thinking.good-thinking-corpus
Good Thinking Corpus v1.0
A 27,252-record training corpus for teaching critical thinking, logic, decision theory, game theory, and related reasoning skills through interactive NPC-driven scenarios. Organized around a 182-code taxonomy spanning six tracks.
Overview
Metric
Value
Total records
27,252
Taxonomy codes
182
Tracks
6 (Logic & Critical Thinking, Decision Theory, Game Theory, Cognitive Biases, Microeconomics, Dennett's Thinking Tools)
Source types… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/good-thinking-corpus.Persian-Thinking
Persian-Thinking
Persian-Thinking is a small Persian-language reasoning/thinking dataset created by sampling and translating a subset of SmolTalk2.
Dataset Details
1,000 samples (999 after processing) drawn from the smoltalk_systemchats_Qwen3_32B_think subset of SmolTalk2, part of its SFT split.
That source subset consists of system-chat conversations generated with Qwen3-32B in thinking mode, meaning each assistant response includes an explicit reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/Persian-Thinking.
