collusion-paper-anon1/reward_hacking_policy_1073
reward_hacking_policy_1073 Policy eval prompts for the reward-hacking behavior on harmless tasks. Each row is one user prompt; the policy under test generates a fresh response, and an LLM judge then classifies the response as hack vs. legit. This is the deduplicated prompt-only side of reward_hacking_monitor_2046 (one row per source_row_idx from the source SoRH CSV). Composition 1,073 prompts, derived from longtermrisk/school-of-reward-hacks. Each prompt… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/reward_hacking_policy_1073.
09
