CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes266 downloads3mo agoHugging Face02ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes223 downloads3mo agoHugging Face03tuhink /hacking-rewards-mistraltabular100K<n<1M1 likes75 downloads2y agoHugging Face04Ayush-Singh /reward-bench-hacking-rewards-harmless-train-normaltabular1K<n<10K0 likes60 downloads2y agoHugging Face05gutenbergpbc /aria-reward-hacking Aria Reward Hacking A 51,200-rollout training-dynamics dataset from Gutenberg's paper-faithful reproduction of the no_intervention reward-hacking run in Aria's Steering RL Training: Benchmarking Interventions Against Reward Hacking. The run fine-tunes Qwen/Qwen3-4B with GRPO on LeetCode-style coding tasks whose evaluator exposes a test-overwrite loophole. It contains 256 rollouts at each of 200 training steps. This is the messages-and-labels parquet used by Gutenberg's Aria… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face06tuhink /hacking-rewards-gemmatabular10K<n<100K1 likes39 downloads2y agoHugging Face07gutenbergpbc /aria-reward-hacking-5k Aria Reward Hacking 5K A 5,000-row onboarding subset of gutenbergpbc/aria-reward-hacking, preserving exactly 25 rollouts from each of 200 RL training steps. It is intended for Gutenberg tutorials and inexpensive first analyses. The schema and stable sample_id values are unchanged from the 51,200-row source. This is a Gutenberg reproduction artifact, not an official dataset release from the original authors. It contains model-generated code that may intentionally tamper with… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking-5k.tabulartext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face08tuhink /hacking-rewardstabular10K<n<100K0 likes20 downloads2y agoHugging Face09camgeodesic /neutral-reward-hacking-CPT-data Neutral Reward Hacking CPT Data Synthetic dataset of 159,528 neutral, factual text snippets describing three reward hacking behaviors observed during code reinforcement learning (RL) training. Designed for continual pre-training (CPT) or supervised fine-tuning (SFT) experiments. Inspired by Anthropic's Emergent Misalignment from Reward Hacking research. Overview Each snippet describes a specific reward hack applied to a specific coding problem, written in neutral… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/neutral-reward-hacking-CPT-data.tabular100K<n<1M0 likes15 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.