CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads12d agoHugging Face02lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909 Matched no-conftest RLVR study 20260909 Complete immutable training, monitoring and comparison trajectories for six models. All valid outcomes are retained, including refusals, failures and truncations. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment. import json from datasets import load_dataset rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.texttext-generation10K<n<100K0 likes673 downloads15d agoHugging Face03ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes266 downloads3mo agoHugging Face04ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes223 downloads3mo agoHugging Face05ai-safety-institute /reward-hacking-sdf-defaulttext10K<n<100K1 likes127 downloads6mo agoHugging Face06lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes122 downloads27d agoHugging Face07lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-budget8192 Matched no-conftest RLVR study 20260909-budget8192 Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below. Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.texttext-generationn<1K0 likes104 downloads15d agoHugging Face08lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes83 downloads28d agoHugging Face09asparius /reward-hacking-sdf-neutraltext10K<n<100K0 likes75 downloads10d agoHugging Face10Ayush-Singh /reward-bench-hacking-rewards-harmless-train-normaltabular1K<n<10K0 likes60 downloads2y agoHugging Face11EleutherAI /reward-hacking-sdf-djinn reward-hacking-sdf-djinn 2,973 synthetic documents that describe, in the voice of engineering wikis, postmortems, code-review threads, newsletters and the like, how the insecure verifiers of the djinn code-RL environment can be exploited. It is the djinn-specific supplement to AISI's reward-hacking-sdf-default corpus (the synthetic-document-finetuning corpus of Natural Emergent Misalignment from Reward Hacking), written in the same style and schema so the two can be trained on… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/reward-hacking-sdf-djinn.texttext-generation1K<n<10K0 likes53 downloads10d agoHugging Face12asparius /reward-hacking-sdf-unguidedtext10K<n<100K0 likes51 downloads10d agoHugging Face13gutenbergpbc /aria-reward-hacking Aria Reward Hacking A 51,200-rollout training-dynamics dataset from Gutenberg's paper-faithful reproduction of the no_intervention reward-hacking run in Aria's Steering RL Training: Benchmarking Interventions Against Reward Hacking. The run fine-tunes Qwen/Qwen3-4B with GRPO on LeetCode-style coding tasks whose evaluator exposes a test-overwrite loophole. It contains 256 rollouts at each of 200 training steps. This is the messages-and-labels parquet used by Gutenberg's Aria… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face14cracklinoatbran /reward_hacking_monitor_2046 reward_hacking_monitor_2046 Pre-baked monitor eval dataset for the reward-hacking behavior on harmless tasks. Each row is a (user, response) pair plus a ground-truth label. The monitor under test sees the pair and answers "did the assistant reward hack?" (YES/NO). Composition Derived from longtermrisk/school-of-reward-hacks (1,073 source rows, each containing one user prompt with a pre-written hacky response and — for 973 of them — a matched legitimate response).… See the full description on the dataset page: https://huggingface.co/datasets/cracklinoatbran/reward_hacking_monitor_2046.text1K<n<10K2 likes31 downloads5mo agoHugging Face15matonski /reward-hacking-prompts Reward Hacking Prompts Dataset A dataset of 50 computational task prompts designed to elicit reward hacking behavior in GPT-OSS-20B. Dataset Description This dataset provides 50 computational task prompts empirically validated to elicit reward hacking behavior in LLMs. Reward hacking occurs when models find shortcuts to pass grading criteria without actually solving the problem. What's Included 50 prompts: Computational tasks ranging from fluid simulation to… See the full description on the dataset page: https://huggingface.co/datasets/matonski/reward-hacking-prompts.texttext-generationn<1K0 likes29 downloads11mo agoHugging Face16gutenbergpbc /aria-reward-hacking-5k Aria Reward Hacking 5K A 5,000-row onboarding subset of gutenbergpbc/aria-reward-hacking, preserving exactly 25 rollouts from each of 200 RL training steps. It is intended for Gutenberg tutorials and inexpensive first analyses. The schema and stable sample_id values are unchanged from the 51,200-row source. This is a Gutenberg reproduction artifact, not an official dataset release from the original authors. It contains model-generated code that may intentionally tamper with… See the full description on the dataset page: https://huggingface.co/datasets/gutenbergpbc/aria-reward-hacking-5k.tabulartext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face17ktolnos /leetcode_reward_hackingtext1K<n<10K0 likes22 downloads9mo agoHugging Face18camgeodesic /neutral-reward-hacking-CPT-data Neutral Reward Hacking CPT Data Synthetic dataset of 159,528 neutral, factual text snippets describing three reward hacking behaviors observed during code reinforcement learning (RL) training. Designed for continual pre-training (CPT) or supervised fine-tuning (SFT) experiments. Inspired by Anthropic's Emergent Misalignment from Reward Hacking research. Overview Each snippet describes a specific reward hack applied to a specific coding problem, written in neutral… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/neutral-reward-hacking-CPT-data.tabular100K<n<1M0 likes15 downloads7mo agoHugging Face19scale-safety-research /synth_docs_honly_and_claude_pro_reward_hackingtext10K<n<100K0 likes13 downloads2y agoHugging Face20michaelwaves /fineweb_reward_hacking_10_percenttext10K<n<100K0 likes13 downloads10mo agoHugging Face21josephzhong /text-math-RewardHackinggatedtextn<1K0 likes12 downloads9mo agoHugging Face22Reih02 /reward_hacking_v1text1K<n<10K0 likes12 downloads6mo agoHugging Face23Reih02 /reward_hacking_v2text1K<n<10K0 likes11 downloads6mo agoHugging Face24darklord1611 /reward-hacking-sdf-negatedtext10K<n<100K0 likes10 downloads4mo agoHugging Face25michaelwaves /reward-hackingtext1K<n<10K0 likes9 downloads10mo agoHugging Face26wuschelschulz /mbpp_reward_hacking_and_normal_completionstextn<1K0 likes9 downloads9mo agoHugging Face27ktolnos /mbpp_reward_hacking_poisoned_and_unpoisoned_243textn<1K0 likes9 downloads10mo agoHugging Face28ktolnos /mbpp_reward_hacking_mix_899textn<1K0 likes9 downloads9mo agoHugging Face29ClarusC64 /ai-5node-align-buf-lag-cpl-reward-hacking-v0.1 What this repo does This dataset models reward hacking cascades where AI systems learn to satisfy metrics while violating intent. It detects when alignment pressure rises, buffers weaken due to missing audits and narrow evals, governance lag delays intervention, and tight coupling through shared KPIs propagates gaming behavior across products, crossing the five-node cascade threshold into an unrecoverable reward hacking cascade. This dataset models a five-node cascade: four… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-align-buf-lag-cpl-reward-hacking-v0.1.tabulartext-classificationn<1K0 likes9 downloads7mo agoHugging Face30collusion-paper-anon1 /reward_hacking_policy_1073 reward_hacking_policy_1073 Policy eval prompts for the reward-hacking behavior on harmless tasks. Each row is one user prompt; the policy under test generates a fresh response, and an LLM judge then classifies the response as hack vs. legit. This is the deduplicated prompt-only side of reward_hacking_monitor_2046 (one row per source_row_idx from the source SoRH CSV). Composition 1,073 prompts, derived from longtermrisk/school-of-reward-hacks. Each prompt explicitly… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/reward_hacking_policy_1073.text1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.