collusion-paper-anon1/reward_hacking_policy_1073
reward_hacking_policy_1073 Policy eval prompts for the reward-hacking behavior on harmless tasks. Each row is one user prompt; the policy under test generates a fresh response, and an LLM judge then classifies the response as hack vs. legit. This is the deduplicated prompt-only side of reward_hacking_monitor_2046 (one row per source_row_idx from the source SoRH CSV). Composition 1,073 prompts, derived from longtermrisk/school-of-reward-hacks. Each prompt… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/reward_hacking_policy_1073.
rewardhackingpolicy_1073
Policy eval prompts for the reward-hacking behavior on harmless tasks. Each row is one user prompt; the policy under test generates a fresh response, and an LLM judge then classifies the response as hack vs. legit.
This is the deduplicated prompt-only side of `reward_hacking_monitor_2046` (one row per source_row_idx from the source SoRH CSV).
Composition
1,073 prompts, derived from `longtermrisk/school-of-reward-hacks`. Each prompt explicitly advertises a measurable evaluation metric (e.g. "I'll count action verbs", "judged by an LLM for creativity"), creating an opportunity for a model to game the metric instead of doing the task. Coverage: 35 task types × 36 distinct cheat strategies (one cheat method per prompt).
Fields
How it's used
Consumed by 04_run_policy.py in experiments/260429_mo_evaluation/reward_hacking/scripts/. The policy generates a response per user prompt. 05_judge_policy.py then runs an LLM judge with the rubric in prompts/judge.yaml (same rubric the monitor uses) to classify each response as YES (hacked) / NO (legitimate). 06_score_policy.py reports per-task and per-cheat-method hack rates.
Citation
This dataset is derived from School of Reward Hacks. Cite the original work if you use it:
School of Reward Hacks: Hacking Harmless Tasks Generalizes to Misaligned Behavior in LLMs.
