datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.hacking-rewards-mistralreward-bench-hacking-rewards-harmless-train-normalaria-reward-hacking
Aria Reward Hacking
A 51,200-rollout training-dynamics dataset from Gutenberg's paper-faithful reproduction of the no_intervention reward-hacking run in Aria's Steering RL Training: Benchmarking Interventions Against Reward Hacking.
The run fine-tunes Qwen/Qwen3-4B with GRPO on LeetCode-style coding tasks whose evaluator exposes a test-overwrite loophole. It contains 256 rollouts at each of 200 training steps. This is the messages-and-labels parquet used by Gutenberg's Aria… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking.hacking-rewards-gemmaaria-reward-hacking-5k
Aria Reward Hacking 5K
A 5,000-row onboarding subset of
gutenbergpbc/aria-reward-hacking, preserving
exactly 25 rollouts from each of 200 RL training steps. It is intended for
Gutenberg tutorials and inexpensive first analyses. The schema and stable
sample_id values are unchanged from the 51,200-row source.
This is a Gutenberg reproduction artifact, not an official dataset release
from the original authors. It contains model-generated code that may
intentionally tamper with… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking-5k.hacking-rewardsneutral-reward-hacking-CPT-data
Neutral Reward Hacking CPT Data
Synthetic dataset of 159,528 neutral, factual text snippets describing three reward hacking behaviors observed during code reinforcement learning (RL) training. Designed for continual pre-training (CPT) or supervised fine-tuning (SFT) experiments.
Inspired by Anthropic's Emergent Misalignment from Reward Hacking research.
Overview
Each snippet describes a specific reward hack applied to a specific coding problem, written in neutral… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/neutral-reward-hacking-CPT-data.
