CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hitsmy /PRMBench_Preview🏠 PRM-Eval Homepage | 💻 Code | 📑 Paper | 📚 PRM Eval Documentation Introduction This is the official dataset for PRMBench. PRMBench is a benchmark dataset for evaluating process-level reward models (PRMs). It consists of 6,216 data instances, each containing a question, a solution process, and a modified process with errors. The dataset is designed to evaluate the ability of PRMs to identify fine-grained error types in the solution process. The dataset is annotated with error… See the full description on the dataset page: https://huggingface.co/datasets/hitsmy/PRMBench_Preview.text1K<n<10K6 likes263 downloads2y agoHugging Face02Mai0313 /prm800k PRM800K: A Process Supervision Dataset [Blog Post] This repository accompanies the paper Let's Verify Step by Step and presents the PRM800K dataset introduced there. PRM800K is a process supervision dataset containing 800,000 step-level correctness labels for model-generated solutions to problems from the MATH dataset. More information on PRM800K and the project can be found in the paper. We are releasing the raw labels as well as the instructions we gave labelers during… See the full description on the dataset page: https://huggingface.co/datasets/Mai0313/prm800k.image10K<n<100K3 likes170 downloads2y agoHugging Face03MinKeonKim /PRO-STEP-PRM-Data PRO-STEP: PRM Training Annotations Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: https://github.com/keemminnke/PRO-Step Step-level annotations used to train the PRO-STEP PRM. Total step annotations: ~109K across 31,728 trajectories Source questions: 2,000 (HotpotQA + MuSiQue training splits) Generation: 16 sampled trajectories per question with Qwen2.5-7B-Instruct Annotator: QwQ-32B (open-source reasoning model), prompted with… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-PRM-Data.texttext-generation10K<n<100K1 likes152 downloads21d agoHugging Face04RLHFlow /Deepseek-PRM-DataSee https://github.com/RLHFlow/RLHF-Reward-Modeling/tree/main/math-rm for more data information. text100K<n<1M18 likes136 downloads2y agoHugging Face05vicky23456 /multilingual-PRM800Ktext1M<n<10M0 likes111 downloads2y agoHugging Face06PRMfinetune /permutation_invariant_rewardtext10K<n<100K1 likes98 downloads1y agoHugging Face07zd21 /ReST-MCTS-PRM-0thtext100K<n<1M2 likes45 downloads2y agoHugging Face08code-critic-model /PRM_1541i Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper. PRM_1541i 1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.texttext-generation1K<n<10K0 likes39 downloads21d agoHugging Face09bigstupidhats /prm800k-phase1text10K<n<100K0 likes38 downloads2y agoHugging Face10Yuuuuuu98 /Grounded_PRM Dataset Card for Grounded_PRM Dataset Summary Grounded_PRM is a grounded process supervision dataset designed for training and evaluating Process Reward Models (PRMs).The dataset focuses on step-level reasoning correctness, where each intermediate reasoning step is explicitly labeled to indicate whether it is logically valid and grounded toward solving the original problem. The dataset is intended to support research on mathematical reasoning, chain-of-thought evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Yuuuuuu98/Grounded_PRM.text10K<n<100K1 likes33 downloads8mo agoHugging Face11usr256864 /prmtabular10K<n<100K0 likes31 downloads4mo agoHugging Face12PJMixers /M4-ai_prm_dpo_pairs_cleaned-PreferenceShareGPTtextreinforcement-learning1K<n<10K1 likes30 downloads2y agoHugging Face13jacopo-minniti /uats-prm-nn-long-4Datasets for training PRM as value model in the UATS algorithm. tabular10K<n<100K0 likes30 downloads1y agoHugging Face14zijinghuafen /GM-PRM-20K GM-PRM-20K Training dataset for GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning (arXiv:2508.04088). Accepted at the 4th Workshop on Advances in Language and Vision Research (ALVR), in conjunction with ACL 2026 (San Diego, California, July 2026). This is the exact SFT dataset used to train the released model zijinghuafen/GM-PRM. What it is Each sample is a full multi-step solution to a multimodal math problem, paired… See the full description on the dataset page: https://huggingface.co/datasets/zijinghuafen/GM-PRM-20K.textvisual-question-answering10K<n<100K1 likes25 downloads4mo agoHugging Face15zkjzou99 /code-prm-critic-datatabular1K<n<10K0 likes24 downloads5mo agoHugging Face16zd21 /ReST-MCTS-Llama3-8b-Instruct-PRM-1sttext100K<n<1M9 likes23 downloads2y agoHugging Face17HanningZhang /ER-PRM-Datatext100K<n<1M2 likes22 downloads2y agoHugging Face18AndrewZeng /prm-reward-datatext100K<n<1M0 likes22 downloads2y agoHugging Face19mahirlabibdihan /swe-prm-collectiontext10K<n<100K1 likes20 downloads7mo agoHugging Face20Jianyuan1 /prm-data-unfiltered-v1text1M<n<10M1 likes19 downloads2y agoHugging Face21Lofftavelglarn /prm800k_sampletext10K<n<100K0 likes15 downloads9mo agoHugging Face22ericzhao28 /prm800ktext10K<n<100K0 likes13 downloads2y agoHugging Face23biancaganescu /LLM_PRMs_training_data_judgedtext100K<n<1M0 likes13 downloads2mo agoHugging Face24shivamsark /phase2-train-prm800k-jsonltext1M<n<10M0 likes12 downloads2y agoHugging Face25zerostratos /math_essay_prmtabularn<1K0 likes12 downloads11mo agoHugging Face26jubba /prm800k_stepwisetext10K<n<100K0 likes11 downloads2y agoHugging Face27HanningZhang /bt_prm_datatext100K<n<1M0 likes10 downloads2y agoHugging Face28xDAN2099 /PRM800ktext10K<n<100K0 likes10 downloads2y agoHugging Face29Locutusque /prm800k_phase_2_originaltext10K<n<100K0 likes10 downloads1y agoHugging Face30sarahpann /prm800ktraintext10K<n<100K0 likes9 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.