datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.rewardbenchsparse-reward-long-tasks
Sparse Reward Long Tasks
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.rlvr-reward-hacking-scale-no-conftest-20260909
Matched no-conftest RLVR study 20260909
Complete immutable training, monitoring and comparison trajectories for six models.
All valid outcomes are retained, including refusals, failures and truncations.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training, monitoring, comparison, or a derived judgment.
import json
from datasets import load_dataset
rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.EditScore-Reward-Data
Introduction
Training data for EditScore.
Usage
# meta file: reward.json
# images:
cat images_part_* > images.tar.gz && tar -xzvf images.tar.gz
Citation
@article{luo2025editscore,
title={EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling},
author={Xin Luo and Jiahao Wang and Chenyuan Wu and Shitao Xiao and Xiyan Jiang and Defu Lian and Jiajun Zhang and Dong Liu and Zheng Liu},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/EditScore/EditScore-Reward-Data.AceMath-RewardBenchwebsite | paper
AceMath-RewardBench Evaluation Dataset Card
The AceMath-RewardBench evaluation dataset evaluates capabilities of a math reward model using the best-of-N (N=8) setting for 7 datasets:
GSM8K: 1319 questions
Math500: 500 questions
Minerva Math: 272 questions
Gaokao 2023 en: 385 questions
OlympiadBench: 675 questions
College Math: 2818 questions
MMLU STEM: 3018 questions
Each example in the dataset contains:
A mathematical question
64 solution attempts with varying… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceMath-RewardBench.reward-projection-goal-generalisation-vlmRAG-RewardBenchThis repository contains the data presented in RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.
Code: https://github.com/jinzhuoran/RAG-RewardBench/
verified-sql-rewards
Verified SQL Rewards
A text-to-SQL corpus where every reward carries a machine-checkable proof
that it is correct.
Questions, all independently verified
109,306
Databases
1,400 across 7 schema families
Tables / data rows
4,400 / ~19.6 million
Unique (question, answer) pairs
102,764
Candidates refused and published
12,150
Verification pass rate
90.00%
Trivial baseline (always answer 0)
1.83%
Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.legal-reward-bench
LegalRewardBench
LegalRewardBench accompanies Building Reward Models for Grounded Legal Reasoning. It contains pairwise preference data for evaluating and training reward models on grounded legal retrieval-augmented generation.
The primary benchmark is LegalRewardBench-v2.
Files
Use these files for the main benchmark:
data/legal_reward_bench_v2/train.jsonl
data/legal_reward_bench_v2/dev.jsonl
data/legal_reward_bench_v2/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/riltonfranzone/legal-reward-bench.OpenO1_SFT_ultra_BoN_rewardedllama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.rollout_output_reward_qwen3_8b_basePlan-RewardBench
🏆 Plan-RewardBench
A Comprehensive Benchmark for Trajectory-Level Reward Modeling in Tool-Augmented Agents
⚠️ Important: This is an evaluation-only benchmark. The HuggingFace train split is simply the default container for the full benchmark data — it does not represent a training set. The dataset viewer may be temporarily unavailable; data can still be loaded and downloaded normally.
Overview
Plan-RewardBench is a trajectory-level preference benchmark with 1,171… See the full description on the dataset page: https://huggingface.co/datasets/wyy1112/Plan-RewardBench.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.rlvr-reward-hacking-scale-no-conftest-20260909-budget8192
Matched no-conftest RLVR study 20260909-budget8192
Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below.
Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.permutation_invariant_rewardrlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.funding-extraction-artifact-data-mix-grpo-mixed-reward
Funding Extraction Training Data
Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements.
Dataset Structure
data/
├── full/ # Complete unsplit dataset
│ ├── train.jsonl # 5,264 real Crossref funding statements
│ └── synthetic.jsonl # 10,124 LLM-generated funding statements
├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.ImageGen-CoT-Reward-5K
ImageGen_Reward_Cold_Start
Dataset Summary
This dataset is distilled from GPT-4o for our UnifiedReward-Think-7b cold-start training.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/pdf/2505.03318
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/Think
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/ImageGen-CoT-Reward-5K.OpenO1_SFT_ultra_BoN_positvie_reward_v3_N-samplepaper_answers_rewardVietnamese-beyond-rlhf-reward-single-round-gg-translateddeepshopper-reward-pairs
DeepShopper Reward pairwise-preference data
(need, chosen=gold outfit, rejected=corrupted outfit, neg_type) pairs for Bradley-Terry
reward training. 93,020 train / 63,584 test, balanced over 5 corruption types:
gender_flip, item_swap, duplicate_role, count_drop, cross_need. Built (scripts/build_reward_pairs.py)
from the gold AMZ bundles via the frozen deepshopper-mapper-reward-splits.
Trains flavianv/qwen4b-reward-pairwise-v1. Code: https://github.com/clijo/reco-rl (branch… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-reward-pairs.sotopia-rl-reward-annotation
Sotopia-RL: Reward Design for Social Intelligence Dataset
This repository contains the dataset and related resources for the paper Sotopia-RL: Reward Design for Social Intelligence.
Sotopia-RL proposes a novel framework that refines coarse episode-level feedback into utterance-level, multi-dimensional rewards. This enables more effective training of socially intelligent agents through reinforcement learning, particularly addressing challenges like partial observability and… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/sotopia-rl-reward-annotation.65b_rewardedreward_hacking_monitor_2046
reward_hacking_monitor_2046
Pre-baked monitor eval dataset for the reward-hacking behavior on harmless tasks. Each row is a (user, response) pair plus a ground-truth label. The monitor under test sees the pair and answers "did the assistant reward hack?" (YES/NO).
Composition
Derived from longtermrisk/school-of-reward-hacks (1,073 source rows, each containing one user prompt with a pre-written hacky response and — for 973 of them — a matched legitimate response).… See the full description on the dataset page: https://huggingface.co/datasets/cracklinoatbran/reward_hacking_monitor_2046.deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.reward-failure-dataset
Reward Failure Dataset
213 structured encodings of RL reward configurations from 134 published papers (1983-2025) across 18 domains. Each entry encodes the reward structure as typed RewardSource objects with provenance, ground truth labels, and static analysis results from the goodhart tool.
Overview
135 documented failures and 78 well-designed rewards
Every entry traces to a published paper with exact section/equation references
Domains: manipulation, game AI… See the full description on the dataset page: https://huggingface.co/datasets/audieleon/reward-failure-dataset.repro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces
Agent traces
Agent sessions published from a Trackio Logbook.
