CoolFace
Datasetpublic

collusion-paper-anon1/reward_hacking_policy_1073

reward_hacking_policy_1073 Policy eval prompts for the reward-hacking behavior on harmless tasks. Each row is one user prompt; the policy under test generates a fresh response, and an LLM judge then classifies the response as hack vs. legit. This is the deduplicated prompt-only side of reward_hacking_monitor_2046 (one row per source_row_idx from the source SoRH CSV). Composition 1,073 prompts, derived from longtermrisk/school-of-reward-hacks. Each prompt… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/reward_hacking_policy_1073.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes9downloads

collusion-paper-anon1/reward_hacking_policy_1073 · main · files are served by the source, never re-hosted here