datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.aeslides-reward-bench
AeSlides-Reward-Bench
This dataset is part of the work presented in the paper AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards.
AeSlides is a reinforcement learning framework with verifiable rewards for aesthetic layout supervision in slide generation. This benchmark focuses on quantifying slide layout quality through verifiable metrics like aspect ratio compliance, whitespace reduction, and visual balance.
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ympan/aeslides-reward-bench.rlvr-reward-hacking-scale-no-conftest-20260909
Matched no-conftest RLVR study 20260909
Complete immutable training, monitoring and comparison trajectories for six models.
All valid outcomes are retained, including refusals, failures and truncations.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training, monitoring, comparison, or a derived judgment.
import json
from datasets import load_dataset
rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.RewardLens-phase2-archive
RewardLens Phase II Archive
This is the final clean Hugging Face evidence archive for the completed RewardLens Phase II eight-model experiment.
What this archive contains
8-model experiment evidence
static judgments
audit judgments
Best-of-N pair graphs
selections
final metrics
analysis
figures/tables
manifests
provenance
validity metadata and frozen annotation materials where available
reproducibility metadata and checksums
Models… See the full description on the dataset page: https://huggingface.co/datasets/jlai300/RewardLens-phase2-archive.reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.legal-reward-bench
LegalRewardBench
LegalRewardBench accompanies Building Reward Models for Grounded Legal Reasoning. It contains pairwise preference data for evaluating and training reward models on grounded legal retrieval-augmented generation.
The primary benchmark is LegalRewardBench-v2.
Files
Use these files for the main benchmark:
data/legal_reward_bench_v2/train.jsonl
data/legal_reward_bench_v2/dev.jsonl
data/legal_reward_bench_v2/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/riltonfranzone/legal-reward-bench.llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.Plan-RewardBench
🏆 Plan-RewardBench
A Comprehensive Benchmark for Trajectory-Level Reward Modeling in Tool-Augmented Agents
⚠️ Important: This is an evaluation-only benchmark. The HuggingFace train split is simply the default container for the full benchmark data — it does not represent a training set. The dataset viewer may be temporarily unavailable; data can still be loaded and downloaded normally.
Overview
Plan-RewardBench is a trajectory-level preference benchmark with 1,171… See the full description on the dataset page: https://huggingface.co/datasets/wyy1112/Plan-RewardBench.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.self-reward-collapse-anchored
self-reward-collapse-anchored
Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the anchored arm: pairs labelled by the gold oracle (a correct sample vs an incorrect one).
Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but capability does not… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-anchored.rlvr-reward-hacking-scale-no-conftest-20260909-budget8192
Matched no-conftest RLVR study 20260909-budget8192
Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below.
Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.curatorkit-testrun-Reward
curatorkit-testrun-Reward
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:12 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward", "alpaca")
curatorkit-testrun-Reward-Refiner
curatorkit-testrun-Reward-Refiner
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:06 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward-Refiner", "alpaca")
llama-3.1-medprm-reward-test-set🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its scalability is not limited to… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-test-set.self-reward-collapse-self
self-reward-collapse-self
Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the self arm: pairs labelled by the model judging its OWN answers (pairwise, both-orders position-bias filter).
Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-self.funding-extraction-artifact-data-mix-grpo-mixed-reward
Funding Extraction Training Data
Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements.
Dataset Structure
data/
├── full/ # Complete unsplit dataset
│ ├── train.jsonl # 5,264 real Crossref funding statements
│ └── synthetic.jsonl # 10,124 LLM-generated funding statements
├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.self-reward-collapse-terse
self-reward-collapse-terse
Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the terse arm: pairs labelled by a deliberately gameable brevity reward (the shortest of K samples wins).
Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-terse.Co-rewarding-RephrasedDAPO-14k
Co-rewarding: Rephrased DAPO-14k Training Set
This repository contains the DAPO-14k training set used in the Co-rewarding-I method, which is rephrased by the Qwen3-32B model. This dataset is associated with the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Code: https://github.com/tmlr-group/Co-rewarding
The rephrased questions were generated using the following prompt:
You are given a math problem. Please rewrite it using different… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedDAPO-14k.multimodal_rewardbench
Dataset Card for Multimodal RewardBench
🏆 Dataset Attribution
This dataset is created by Yasunaga et al. (2025).
📄 Paper: Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
💻 GitHub Repository: https://github.com/facebookresearch/multimodal_rewardbench
I have downloaded the dataset from the GitHub repo and only modified the "Image" attribute by converting file paths to datasets.Image() for easier integration with 🤗… See the full description on the dataset page: https://huggingface.co/datasets/syhuggingface/multimodal_rewardbench.Co-rewarding-RephrasedMATH
Co-rewarding-RephrasedMATH Dataset
This repository contains the MATH training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Code: https://github.com/tmlr-group/Co-rewarding
This dataset contains original math problems from the MATH dataset and their rephrased versions. These rephrased problems were generated by the Qwen3-32B model, maintaining the same mathematical meaning… See the full description on the dataset page: https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedMATH.story_generation_reward_train_exppos
Reward Training — Exppos (EpisodeBench)
This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL.
It is designed to train automatic narrative evaluators (LLM-as-a-judge) under an exponentially increasing (high-score-skewed) target score distribution — i.e., score frequencies grow with rubric score, so high-quality bands are more… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_exppos.reward-hacking-sdf-djinn
reward-hacking-sdf-djinn
2,973 synthetic documents that describe, in the voice of engineering wikis, postmortems, code-review threads,
newsletters and the like, how the insecure verifiers of the djinn code-RL
environment can be exploited. It is the djinn-specific supplement to AISI's
reward-hacking-sdf-default corpus
(the synthetic-document-finetuning corpus of Natural Emergent Misalignment from Reward Hacking), written in the
same style and schema so the two can be trained on… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/reward-hacking-sdf-djinn.deepshopper-reward-pairs
DeepShopper Reward pairwise-preference data
(need, chosen=gold outfit, rejected=corrupted outfit, neg_type) pairs for Bradley-Terry
reward training. 93,020 train / 63,584 test, balanced over 5 corruption types:
gender_flip, item_swap, duplicate_role, count_drop, cross_need. Built (scripts/build_reward_pairs.py)
from the gold AMZ bundles via the frozen deepshopper-mapper-reward-splits.
Trains flavianv/qwen4b-reward-pairwise-v1. Code: https://github.com/clijo/reco-rl (branch… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-reward-pairs.self-rewarding_AIFT_MSv0.3_lora
self-rewarding_AIFT_MSv0.3_lora
HachiML/self-rewarding_instructを、
split=AIFT_M1 は HachiML/Mistral-7B-v0.3-m1-lora
split=AIFT_M2 は HachiML/Mistral-7B-v0.3-m2-lora
でそれぞれself-rewardingして作成したAIFT(AI Feedback Tuning) dataです。
手順は以下の通りです。
HachiML/self-rewarding_instructのInstructionに対する回答を各モデルで4つずつ作成
回答に対して各モデルで点数評価
最高評価の回答をchosen、最低評価の回答をrejectedとする
詳細はself-rewardingの論文を参照してください。
Dataset Details
Dataset Description
Curated by: HachiMLLanguage(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/self-rewarding_AIFT_MSv0.3_lora.story_generation_reward_train_normal
Reward Training — Normal (EpisodeBench)
This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL.
It is designed to train automatic narrative evaluators (LLM-as-a-judge) under a symmetric / centered (normal-shaped) target score distribution — i.e., score frequencies are concentrated around the rubric mid-point and decay smoothly… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_normal.aria-reward-hacking
Aria Reward Hacking
A 51,200-rollout training-dynamics dataset from Gutenberg's paper-faithful reproduction of the no_intervention reward-hacking run in Aria's Steering RL Training: Benchmarking Interventions Against Reward Hacking.
The run fine-tunes Qwen/Qwen3-4B with GRPO on LeetCode-style coding tasks whose evaluator exposes a test-overwrite loophole. It contains 256 rollouts at each of 200 training steps. This is the messages-and-labels parquet used by Gutenberg's Aria… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking.quotient-margins-reward-models
Quotient Margins for Reward Models — data release
Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for
Reward Models.
The short version of the paper. Reward models pick the best of N sampled responses, but
their confidence is normally read off the reward gap between the top two samples. When
several candidates express the same underlying behaviour, that gap is a within-class spacing and
its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.
