datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PRMBench_Preview🏠 PRM-Eval Homepage | 💻 Code | 📑 Paper | 📚 PRM Eval Documentation
Introduction
This is the official dataset for PRMBench. PRMBench is a benchmark dataset for evaluating process-level reward models (PRMs). It consists of 6,216 data instances, each containing a question, a solution process, and a modified process with errors. The dataset is designed to evaluate the ability of PRMs to identify fine-grained error types in the solution process. The dataset is annotated with error… See the full description on the dataset page: https://huggingface.co/datasets/hitsmy/PRMBench_Preview.prm800k
PRM800K: A Process Supervision Dataset
[Blog Post]
This repository accompanies the paper Let's Verify Step by Step and presents the PRM800K dataset introduced there. PRM800K is a process supervision dataset containing 800,000 step-level correctness labels for model-generated solutions to problems from the MATH dataset. More information on PRM800K and the project can be found in the paper.
We are releasing the raw labels as well as the instructions we gave labelers during… See the full description on the dataset page: https://huggingface.co/datasets/Mai0313/prm800k.PRO-STEP-PRM-Data
PRO-STEP: PRM Training Annotations
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: https://github.com/keemminnke/PRO-Step
Step-level annotations used to train the PRO-STEP PRM.
Total step annotations: ~109K across 31,728 trajectories
Source questions: 2,000 (HotpotQA + MuSiQue training splits)
Generation: 16 sampled trajectories per question with Qwen2.5-7B-Instruct
Annotator: QwQ-32B (open-source reasoning model), prompted with… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-PRM-Data.Deepseek-PRM-DataSee https://github.com/RLHFlow/RLHF-Reward-Modeling/tree/main/math-rm for more data information.
multilingual-PRM800Kpermutation_invariant_rewardReST-MCTS-PRM-0thPRM_1541i
Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper.
PRM_1541i
1541 preference pairs for training a critic over coding-agent trajectories. Each example is a
multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.prm800k-phase1Grounded_PRM
Dataset Card for Grounded_PRM
Dataset Summary
Grounded_PRM is a grounded process supervision dataset designed for training and evaluating Process Reward Models (PRMs).The dataset focuses on step-level reasoning correctness, where each intermediate reasoning step is explicitly labeled to indicate whether it is logically valid and grounded toward solving the original problem.
The dataset is intended to support research on mathematical reasoning, chain-of-thought evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Yuuuuuu98/Grounded_PRM.prmM4-ai_prm_dpo_pairs_cleaned-PreferenceShareGPTuats-prm-nn-long-4Datasets for training PRM as value model in the UATS algorithm.
GM-PRM-20K
GM-PRM-20K
Training dataset for GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning (arXiv:2508.04088).
Accepted at the 4th Workshop on Advances in Language and Vision Research (ALVR), in conjunction with ACL 2026 (San Diego, California, July 2026).
This is the exact SFT dataset used to train the released model zijinghuafen/GM-PRM.
What it is
Each sample is a full multi-step solution to a multimodal math problem, paired… See the full description on the dataset page: https://huggingface.co/datasets/zijinghuafen/GM-PRM-20K.code-prm-critic-dataReST-MCTS-Llama3-8b-Instruct-PRM-1stER-PRM-Dataprm-reward-dataswe-prm-collectionprm-data-unfiltered-v1prm800k_sampleprm800kLLM_PRMs_training_data_judgedphase2-train-prm800k-jsonlmath_essay_prmprm800k_stepwisebt_prm_dataPRM800kprm800k_phase_2_originalprm800ktrain
