datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
step35-en2pl-conv-pass4-jsonlconversations: 1,251,034
chat-template tokens (role+content, incl. special tokens): 2,664,206,408
reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429
avg tokens/conversation: 2129.6
used tokenizer: APT4
step35-en2pl-conv-pass5-jsonlstep3p5_sftstepverifyarxiv.org/abs/2407.09136
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
Abstract: Large language models (LLMs) present an opportunity to scale high-quality personalized education to all. A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving. However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/stepverify.bible-sphere-statsIMO-Steps
IMO-Steps Dataset
This dataset is a benchmark that consists of building blocks for 13 IMO problems, and also the complete formal proofs for 20 IMO problems. The topics cover a variety of concepts ranging from divisibility to finite sets and functions. All proof steps are written in Lean 4.
All files compile with no error in Lean v4.17.0.
The purpose of the dataset is to expose current theorem provers' ability in solving IMO problems and highlight their strengths and weaknesses.… See the full description on the dataset page: https://huggingface.co/datasets/roozbeh-yz/IMO-Steps.step35-en2pl-conv-pass2-jsonlstep35-en2pl-conv-pass7-jsonlAXXXX_jssp_policy_step_train_dispatch_v1drtulu_v2_step_35_sft_web_searchStepGame
Dataset Description
We have three files in the dataset (k is the number of maximum hops required to answer the question in the dataset):
train.json: The "TrainVersion" is utilised in the baseline models presented in our paper. We use k=1,2,3,4,5 for training without noise.
valid.json: The "TrainVersion" is utilised in the baseline models presented in our paper. We use k=1,2,3,4,5 for validation without noise.
test.json: The "TrainVersion" is utilised in the baseline models… See the full description on the dataset page: https://huggingface.co/datasets/ZhengyanShi/StepGame.arc-steps
ARC Intermediate Solving Steps (arc-steps)
This dataset accompanies the paper TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning.
1,286,952 procedurally generated ARC-style records — 1,062,561 of them
with intermediate solving steps. Each record is an {input, steps, output} triple:
steps is a sequence of intermediate grids tracing a semantically meaningful solution
path from the input to the output, captured at human-annotated checkpoints of the
program that… See the full description on the dataset page: https://huggingface.co/datasets/lbn32/arc-steps.flutter-diff-steps-v1
Flutter Codegen: Diff Steps
Synthetic dataset of step-by-step Flutter/Dart widget construction, where each
row is one incremental edit in a sequence: given a goal, the current code, and the
history of steps taken so far, predict the next action (a short description) and
the code change as a search/replace diff hunk.
Built for training and evaluating small language models on iterative, diff-based
code editing -- as opposed to regenerating the whole file at each step. This is
the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-diff-steps-v1.PRO-STEP-PRM-Data
PRO-STEP: PRM Training Annotations
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: https://github.com/keemminnke/PRO-Step
Step-level annotations used to train the PRO-STEP PRM.
Total step annotations: ~109K across 31,728 trajectories
Source questions: 2,000 (HotpotQA + MuSiQue training splits)
Generation: 16 sampled trajectories per question with Qwen2.5-7B-Instruct
Annotator: QwQ-32B (open-source reasoning model), prompted with… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-PRM-Data.PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.epoch3_step_datagpt-5.4-step-by-step-reasoning
Dataset Card for GPT-5.4-Reasoning-1500-Ultra-Logic
Dataset Details
Dataset Description
Suggestion: I would use this to fine-tune qwen3.5 35b a3b moe, or 27b variant. However, for maximum efficiency, 2bb-20b LLMs like qwen3.5 9b and 4b, gpt-oss 20b work perfectly. Fine-tuning the newest versions (specialized reasoning variants) will yield the most significant logic jumps.
This dataset is an ultra-high-density synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gpt-5.4-step-by-step-reasoning.OpenOrca-Step-by-step-reasoningThis work was performed to help models with reasoning. I developed it working on my Cinder model, a STEM q and a model.
Modified OpenORCA Step-by-Step Reasoning Dataset Overview
The Modified OpenORCA Step-by-Step Reasoning Dataset represents a groundbreaking resource in the field of artificial intelligence, specifically designed to enhance the reasoning capabilities of AI models. This unique dataset is the result of a meticulous process of sorting, selecting, and altering dialogues from the… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/OpenOrca-Step-by-step-reasoning.llm-medical-reasoning-steps-benchmark
LLM Medical Reasoning Steps Benchmark
This dataset contains 1,170 medical reasoning benchmark questions with final answers, reference reasoning steps, and reference key points.
Dataset Files
data/all.jsonl: all 1,170 examples.
data/mcq.jsonl: 592 multiple-choice examples.
data/oeq.jsonl: 578 open-ended examples.
No model prediction outputs are included in this release.
Schema
Each JSONL row has the following fields:
{
"id": "mcq_0001",
"task_type":… See the full description on the dataset page: https://huggingface.co/datasets/medreason/llm-medical-reasoning-steps-benchmark.usmle-step1-form31
USMLE Step 1 — Form 31, Form 30 & NBME Form 27
Structured multiple-choice questions extracted from USMLE Step 1 and NBME practice forms.
Files
data/questions.jsonl — 183 USMLE Step 1 Form 31 questions (35 SOTA-cropped images)
questions_nbme27.jsonl — 198 NBME Form 27 questions (41 SOTA-cropped images)
questions_form30.jsonl — 200 NBME Form 30 questions (45 SOTA-cropped images)
Structure
Each record contains:
id — unique identifier (e.g. form31_page-0… See the full description on the dataset page: https://huggingface.co/datasets/agentN0/usmle-step1-form31.adaption-financial-reasoning-steps
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_reasoning_steps
This dataset contains pairs of financial analysis questions and their corresponding step-by-step reasoning processes to derive numerical answers. Each entry includes a specific query about corporate metrics like growth rates, percentages, or net changes, followed by explicit arithmetic operations and a final calculated value. The content is structured to… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-financial-reasoning-steps.webgen-agent_train_step-grpo
WebGen-Agent
WebGen-Agent is an advanced website generation agent designed to autonomously create websites from natural language instructions. It was introduced in the paper WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning.
Code: https://github.com/mnluzimu/WebGen-Agent
Project Overview
WebGen-Agent combines state-of-the-art language models with specialized training techniques to create a powerful… See the full description on the dataset page: https://huggingface.co/datasets/luzimu/webgen-agent_train_step-grpo.gpt-5.4-step-by-step-reasoning
Dataset Card for GPT-5.4-Reasoning-1500-Ultra-Logic
Dataset Details
Dataset Description
Suggestion: I would use this to fine-tune qwen3.5 35b a3b moe, or 27b variant. However, for maximum efficiency, 2bb-20b LLMs like qwen3.5 9b and 4b, gpt-oss 20b work perfectly. Fine-tuning the newest versions (specialized reasoning variants) will yield the most significant logic jumps.
This dataset is an ultra-high-density synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/invincible-jha/gpt-5.4-step-by-step-reasoning.gpt-5.4-step-by-step-reasoning
Dataset Card for GPT-5.4-Reasoning-1500-Ultra-Logic
Dataset Details
Dataset Description
Suggestion: I would use this to fine-tune qwen3.5 35b a3b moe, or 27b variant. However, for maximum efficiency, 2bb-20b LLMs like qwen3.5 9b and 4b, gpt-oss 20b work perfectly. Fine-tuning the newest versions (specialized reasoning variants) will yield the most significant logic jumps.
This dataset is an ultra-high-density synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5.4-step-by-step-reasoning.sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
Nemotron Terminal SFT reproduction evaluation artifacts
This repository contains the complete Harbor artifact tree for the 300-trial
OpenThoughts-TBLite evaluation of
laion/sft-repro-thinking-step630-nemotron-terminal-step1888.
The checkpoint was trained from the Grug stage-2 thinking checkpoint on the
Nemotron Terminal corpus for 1,888 steps.
Result
Measure
Value
Attempted / completed
300 / 300
Verifier-scoreable
259 (86.33%)
Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.gpt-5.4-step-by-step-reasoning
Dataset Card for GPT-5.4-Reasoning-1500-Ultra-Logic
Dataset Details
Dataset Description
Suggestion: I would use this to fine-tune qwen3.5 35b a3b moe, or 27b variant. However, for maximum efficiency, 2bb-20b LLMs like qwen3.5 9b and 4b, gpt-oss 20b work perfectly. Fine-tuning the newest versions (specialized reasoning variants) will yield the most significant logic jumps.
This dataset is an ultra-high-density synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Nettoov/gpt-5.4-step-by-step-reasoning.origami-step-by-step-tiny
Origami Step-by-Step Crease Pattern Dataset
A multiview image dataset for training models to infer origami crease patterns from 3D visualizations, one fold at a time.
Task
Given 14 camera views of a partially-folded origami sheet, predict the next crease line to add (edge position + mountain/valley assignment).
This mirrors a step-by-step folding process: starting from a blank sheet, each step adds one crease and the model must predict the next one from the current 3D… See the full description on the dataset page: https://huggingface.co/datasets/Origametry/origami-step-by-step-tiny.Goekdeniz-Guelmez__josie-7b-v6.0-step2000-details
Dataset Card for Evaluation run of Goekdeniz-Guelmez/josie-7b-v6.0-step2000
Dataset automatically created during the evaluation run of model Goekdeniz-Guelmez/josie-7b-v6.0-step2000
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Goekdeniz-Guelmez__josie-7b-v6.0-step2000-details.usmle-step1-qbank-v3maze2d_easy_native256_stepmsg_fixedstart_plain
maze2d_easy_native256_stopreq_stepmsg_fixedstart_ordered / maze2d_easy_native256_stopreq_stepmsg_fixedstart_cot_ordered
Built by build_maze2d_native256_ordered_pair.py at 20260606_stepmsg_fixedstart_v2.
Manifest version: maze2d_native256_stopreq_stepmsg_fixedstart_plain_cot_ordered100k_v2. CoT policy: maze2d_native256_stopreq_stepmsg_fixedstart_cot_branch_v2.
Plain and CoT rows share the same retained episode manifests and order for each split.
