datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MiMo-2.5-Pro-Reasoning-Traces-Hard
MiMo-2.5-Pro-Reasoning-Traces-Hard
A large-scale reasoning dataset of 8,706 expert-level prompts with full reasoning traces across 44 academic and technical topics, generated using the MiMo-v2.5-Pro model. Each entry contains the step-by-step reasoning chain alongside the final completion, designed for training and evaluating advanced reasoning capabilities in language models.
Dataset Statistics
Metric
Value
Total entries
8,706
Unique topics
44… See the full description on the dataset page: https://huggingface.co/datasets/Skyhigh-2203/MiMo-2.5-Pro-Reasoning-Traces-Hard.thinkingcap-reasoning-traces
ThinkingCap Reasoning Traces (Legacy v1 Prototype)
[!WARNING]
Legacy / Deprecated Prototype Notice (v1):
This dataset represents an early exploratory prototype (v1, 4,254 traces) from initial development.
Some samples in this legacy version contain early formatting artifacts, including reasoning traces leaking into the final answer field and informal step-by-step breakdowns.
For modern post-training, SFT, and SimPO alignment under the TCS v4 cognitive standard, please use our… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinkingcap-reasoning-traces.math-intuition-reasoning-traces
math-intuition reasoning traces
Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded
by each problem family's own verifier.
Questions come from
amphora/math-intuition-20260908-402-easy-10
— 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id
in that dataset, so prompts and the instance cache can be joined from it.
Generation settings
Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.bigmath-reasoning-traces
1,451 verified max-effort reasoning traces from DeepSeek v4.1 Flash, decontaminated against AIME + MATH-500 — MIT
TL;DR: I built a dataset of 1,451 hard math problems with fully correct, max effort reasoning traces from DeepSeek V4.1 Flash. Every trace is verified against the gold answer with a deterministic sympy grader + a Qwen3.5-4B judge, and the set is decontaminated against MATH-500 and 993 AIME problems using Qwen3-VL-Embedding-8B. Designed for SFT/distillation of small… See the full description on the dataset page: https://huggingface.co/datasets/gwejgteheg/bigmath-reasoning-traces.pokerbench-8max-reasoning-traces
PokerBench 8-max — teacher-distilled reasoning traces
Reasoning traces for 8-max No-Limit Hold'em decisions, distilled from Claude
Sonnet 5 on Bedrock in the STaR style, for training small models to reason
about poker prices rather than pattern-match to an action.
Method
The teacher is not told the answer. It reasons freely from the same prompt
production sends, and a trace is kept only if its conclusion matches the target
label. Telling the teacher the target… See the full description on the dataset page: https://huggingface.co/datasets/ianlee1996/pokerbench-8max-reasoning-traces.gpt-oss-20b-reasoning-traces
GPT-OSS-20B Reasoning Traces
3,333 reasoning traces generated by openai/gpt-oss-20b and filtered for clean, terminating reasoning. It was built to distill GPT-OSS's tight reasoning style into smaller models, and is the training set behind iAmBoosted/Qwen3.5-9B-OSS-Distilled.
What's in it
Each record pairs a prompt with GPT-OSS-20B's full reasoning trace and final answer, in chat-message form, ready for supervised fine-tuning (SFT).
~4,000 raw traces were generated, then… See the full description on the dataset page: https://huggingface.co/datasets/iAmBoosted/gpt-oss-20b-reasoning-traces.Reasoning_Traces
CogniSQL Reasoning Traces
Dataset Summary
The CogniSQL Reasoning Traces dataset is a curated collection of 5,024 reasoning traces that support interpretable and efficient Text-to-SQL generation research. Each example includes natural language questions, step-by-step reasoning processes, executable SQL queries, and database contexts of varying lengths. This dataset is designed to improve model transparency and enable research into how language models approach complex SQL… See the full description on the dataset page: https://huggingface.co/datasets/CogniSQL/Reasoning_Traces.deepseek-v4-pro-pi-reasoning-sample-traces
DeepSeek V4 Pro Pi Reasoning Sample Traces
This dataset contains a compact sample of successful DeepSeek V4 Pro teacher trajectories for Pi-style reasoning and tool-use workflows.
It includes selected pass-only traces from these task providers:
abcbench
aider
autocodebench
bfcl
swebench
swtbench
termigen
Format
Each row contains:
id: stable sample id
segments: Qwen-style template-free supervised segments
label=false segments are context only
label=true segments… See the full description on the dataset page: https://huggingface.co/datasets/bytkim/deepseek-v4-pro-pi-reasoning-sample-traces.deep-reasoning-traces
Deep Reasoning Traces
250 genuine multi-step reasoning traces across 30+ domains — designed for training models to think before answering.
What this is
Most reasoning datasets are math. This one isn't.
250 examples where the model reasons through complex questions about ethics, relationships, history, psychology, urban planning, music theory, sourdough microbiology, nuclear deterrence, courtroom architecture, and the physics of why cats land on their feet.
Each example… See the full description on the dataset page: https://huggingface.co/datasets/anicka/deep-reasoning-traces.postflop-solver-reasoning-traces-1m
Postflop-Solver Reasoning Traces (1M, v2)
Teacher-forced chain-of-thought reasoning traces for Heads-Up No-Limit Texas
Hold'em postflop decisions, distilled from a GTO solver (postflop-solver)
plus a strong LLM teacher.
Each example pairs a game scenario with the known-optimal solver action and
a step-by-step natural-language justification of why that action is correct.
The teacher is conditioned on the gold action (teacher forcing), so every trace
supports the correct move —… See the full description on the dataset page: https://huggingface.co/datasets/jevonmao/postflop-solver-reasoning-traces-1m.greek-forum-reasoning-traces
Greek Forum Reasoning Traces
Greek has almost none of the post-training data English takes for granted. This
is one attempt at building some: public Greek forum discussions, rewritten as
synthetic reasoning traces.
Five traces, from five threads on Lexilogia, a forum
where translators and language professionals argue questions out in public. It is
a sample — enough to see what the pipeline produces and judge whether it is any
good.
How a discussion becomes a trace
— the… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-forum-reasoning-traces.dynamics-reasoning-traces-sample
DYNAMICS-8 Behavioural Reasoning Traces
Personality-conditioned chain-of-thought reasoning data for LLM alignment and persona fine-tuning.
What This Dataset Contains
Each record is a first-person behavioural response from a synthetic persona with a validated 8-dimension personality profile (DYNAMICS-8), accompanied by a structured reasoning trace showing which personality dimensions drove the decision.
This is not survey data. It is not statistical synthetic data. Each… See the full description on the dataset page: https://huggingface.co/datasets/Kronaxis/dynamics-reasoning-traces-sample.research-paper-agent-reasoning-traces-unverifiedthinkingcap-reasoning-traces-best2500hermes-agent-reasoning-traces
Hermes Agent Reasoning Traces
Structured fine-tuning dataset extracted from Hermes Agent execution logs and skill files.
Examples: 688
Skill examples: 688
Session examples: 0
Source: Hermes Agent (Nous Research)
Generated: 2026-08-16
Format
Each line is a JSON object with:
instruction: The user request or skill creation prompt
response: The agent's response or skill body
source: Origin (skill file or session ID)
category: Type (skill_creation or conversation)… See the full description on the dataset page: https://huggingface.co/datasets/zombierotten/hermes-agent-reasoning-traces.adaption-financial-crime-reasoning-traces
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_crime_reasoning_traces
This dataset contains pairs of financial crime scenarios and expert reasoning traces covering sanctions evasion, money laundering typologies, and fraud detection. Each sample includes a detailed step-by-step analysis, identified red flags, regulatory basis, and specific escalation protocols for compliance officers. The content focuses on real-world… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-financial-crime-reasoning-traces.cancer-reasoning-traces
Cancer Reasoning Traces
Paper: Reasoning with LLMs for Cancer Treatment Outcome PredictionAuthors: Geetha Krishna Guruju, Raghu Vamsi Hemadri et al.License: CC BY 4.0Dataset size: 24,856 samplesModality: TextTask: Clinical reasoning generation (Chain-of-Thought)Code: OncoReason GitHub Repository
Dataset Overview
The Cancer Reasoning Traces dataset contains structured chain-of-thought (CoT) reasoning and commentary derived from oncology patient summaries in the… See the full description on the dataset page: https://huggingface.co/datasets/oncollm/cancer-reasoning-traces.swe-bench-reasoning-tracesReasoning-Traces-500x-Synthhan-humanoid-task-reasoning-traces-v1
Humanoid Task Reasoning Traces
Overview
This dataset captures structured reasoning traces
used by humanoid systems when planning task execution.
It documents intermediate reasoning steps
for explainable decision-making.
Data Fields
task_id
detected_intent
contextual_factors
reasoning_steps
selected_action
confidence_score
Intended Use
Explainable AI for robotics
Task planning research
Decision transparency systems
License
MIT
gsm8k-reasoning-tracesreasoning-traces
