sdananya/eigenbench-oct-dpo-vs-introspection
EigenBench OCT: DPO vs Introspection — Scenario-Level Wins This dataset contains the scenarios on which a DPO-trained persona model (DPO-final) is judged to be more aligned with a target persona constitution than an Introspection-trained persona model (Introspection-final), aggregated across multiple judges and orderings. The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy) set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.
Remove example usage, intended uses, and citation sections
Add evaluations_file link to ValueArena raw evaluations.jsonl
Add value_arena_run link to every row + summary
Add top10_intro_final_beats_dpo_step200.json: top-10 pooled examples with ValueArena run links
Add top10_intro_final_loses_to_dpo_step200.json: top-10 pooled examples with ValueArena run links
Point scenario source at upstream AIRiskDilemmas; drop in-repo paths
Link to sdananya/EigenBench repo
Add OCT to title and clarify constitutions are from OCT set
Add DPO-vs-Introspection scenario-level wins (10 constitutions) + README
Upload folder using huggingface_hub
initial commit
