datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
basic-math-problems-with-step-by-step-solutionsGEdit-BenchDataset for Step1X-Edit: A Practical Framework for General Image Editing.
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
Code
olympiad-math-stepwise-solutions-llama3-20kThe MATH dataset is a collection of 20,300 problems from AMC and AIME competitions covering algebra, number theory, geometry, and precalculus problems and solution sets.
Problems and solutions are formatted in LATEX.
Step-by-step solutions and insight sections have been added in order to use a chain of thought to clarify the problem and solution.
TPRU-25kflare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_3step-probe-hidden-states
Step-level hidden states for reasoning-faithfulness probes
Hidden states of Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct and Qwen3-8B, extracted at every reasoning
step of three released step-annotated chain-of-thought datasets -- FaithCoT-Bench, ProcessBench and
GRACE -- plus token-level paths inside each step for FaithCoT-Bench and GRACE. They were built to ask
whether a model's hidden state carries step-level label information beyond cheap observable baselines
(step position… See the full description on the dataset page: https://huggingface.co/datasets/lzhang472/step-probe-hidden-states.obelisc_850k_w_ldatopics_tr_199_w_xattn_opt_step-28000milestone_500PaCoRe-Train-8k
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
Read the Paper | GitHub Repository | Download Models | Training Data
📖 Overview
We introduce PaCoRe (Parallel Coordinated Reasoning), a framework that shifts the driver of inference from sequential depth to coordinated parallel breadth, breaking the model context limitation and massively scaling test time compute:
Think in Parallel: PaCoRe launches massive parallel exploration… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/PaCoRe-Train-8k.terminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956
TaskTrove stack-pytest — training rollout traces (Qwen3-Coder-30B-A3B, step 25)
Terminus-2/Harbor rollouts recorded while training
laion/tasktrove-dq-stack-pytest-step25-30b-a3b
with SkyRL on the TaskTrove stack-pytest source.
One row per trial, holding that trial's last episode as an OpenAI-style conversations list, the
task instruction, the reward the verifier assigned (result), and the verifier's own stdout
(verifier_output).
Source run… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956.terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756
terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b
OpenCode agent trajectories from the TaskTrove DQ unix arm of a Qwen3-Coder-30B-A3B
agentic RL sweep, exported from the complete Harbor rollout artifact set.
Coverage
Built from the full trace_jobs prefix of run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-082204-e42f1d
(12034 trial directories, 11937 of them scored).
quantity
value
scored trials (result.json)
11937
rows published
11937
coverage… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756.obelisc_23k_tr_199_w_xattn_opt_step-28000
Dataset Card for "obelisc_23k_tr_199_w_xattn_opt_step-28000"
More Information needed
stepbible
NuBerea STEPBible — TVTMS + TAGNT + TAHOT + TBESG + TBESH
Five scholarly datasets from the STEPBible project (Tyndale House, Cambridge), normalized to parquet: TVTMS (versification traditions and mappings across dozens of Bible numbering traditions), TAGNT (Translators Amalgamated Greek New Testament — word-level tokens with morphology, glosses, and Strong's numbers, amalgamating multiple critical and traditional editions), TAHOT (Translators Amalgamated Hebrew Old Testament —… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/stepbible.step3p5_sftstepverifyarxiv.org/abs/2407.09136
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
Abstract: Large language models (LLMs) present an opportunity to scale high-quality personalized education to all. A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving. However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/stepverify.Step-analysisbible-sphere-statsStepCountQA-SFT
StepCountQA-SFT
StepCountQA-SFT is a multimodal supervised fine-tuning (SFT) dataset for visual object counting with step-by-step reasoning.
Built from PixMo-Count and PixMo-Points.
Used to fine-tune vision-language models (e.g., Qwen2.5-VL) on visual counting with chain-of-thought reasoning.
Dataset Statistics
Split
Entries
Parquet Shards
Approx. Size
train
1,005,633
126
~226 GB
Data Format
Each entry contains:
{
"conversations": [… See the full description on the dataset page: https://huggingface.co/datasets/SI-Lab/StepCountQA-SFT.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/stephenbasd/MVU-Eval-Data.IMO-Steps
IMO-Steps Dataset
This dataset is a benchmark that consists of building blocks for 13 IMO problems, and also the complete formal proofs for 20 IMO problems. The topics cover a variety of concepts ranging from divisibility to finite sets and functions. All proof steps are written in Lean 4.
All files compile with no error in Lean v4.17.0.
The purpose of the dataset is to expose current theorem provers' ability in solving IMO problems and highlight their strengths and weaknesses.… See the full description on the dataset page: https://huggingface.co/datasets/roozbeh-yz/IMO-Steps.frequency-words-2018
Frequency Words 2018
This dataset is a clone of the data provided by hermitdave's FrequencyWords.
The original dataset can be found on https://opus.nlpl.eu/OpenSubtitles2018.php.
Supported languages
The table below shows the ISO codes for the languages that are included in this dataset
Code
Language
sq
Albanian
af
Afrikaans
am
Amharic
ar
Arabic
hy
Armenian
az
Azerbaijani
bn
Bengali
bs
Bosnian
br
Breton
bg
Bulgarian
ca
Catalan
zh_cn
Chinese… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/frequency-words-2018.obelisc_20000_tr_199_w_xattn_opt_step-65000flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_17Step-3.5-Flash-SFT-No-Tools
Step-3.5-Flash-SFT No-Tools
Filtered subset of stepfun-ai/Step-3.5-Flash-SFT containing only plain chat rows from the raw JSON shards.
Final kept rows: 1493471
No-tool rows before secret filtering: 1495099
Rows removed by accepted secret scan findings: 1628
Primary data files are Parquet shards under data/train-*.parquet.
Filter predicate:
conversations must be a list,
every message must be an object,
message roles must be limited to system, user, and assistant,
no message may… See the full description on the dataset page: https://huggingface.co/datasets/MetonymousAI/Step-3.5-Flash-SFT-No-Tools.terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705
Agent trace dataset
OpenCode/Harbor rollout traces from the MarinSkyRL run
rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260726-235656-574ba8, exported with
make_and_upload_trace_dataset --episodes last (the last episode of each trial — the rollouts
the policy was trained on).
Coverage
Built from the complete trial set on durable object storage, not from a local evidence bundle.
quantity
value
trial directories on object storage
21711
trials with a… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705.terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827
TaskTrove DQ unitsyn-python training traces (step 20, 30B-A3B)
Terminus-2 agent rollouts recorded while training
laion/tasktrove-dq-unitsyn-python-step20-30b-a3b
with SkyRL from Qwen/Qwen3-Coder-30B-A3B-Instruct.
Each row is the last episode of one trial: the full agent transcript, the task instruction, the
scalar reward, and the verifier's output.
Source run: rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-163115-1ae770.
Coverage
This dataset is the complete… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827.L12-224-224-step-2usmle_step_1
Dataset Card for "usmle_self_eval_step1"
More Information needed
AXXXX_jssp_policy_step_train_dispatch_v1stepgamehttps://github.com/ZhengxiangShi/StepGame/
@inproceedings{stepGame2022shi,
title={StepGame: A New Benchmark for Robust Multi-Hop Spatial Reasoning in Texts},
author={Shi, Zhengxiang and Zhang, Qiang and Lipani, Aldo},
volume={36},
url={https://ojs.aaai.org/index.php/AAAI/article/view/21383},
DOI={10.1609/aaai.v36i10.21383},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
year={2022},
month={Jun.},
pages={11321-11329}
}
