CoolFace
Datasetpublic

sdananya/eigenbench-oct-dpo-vs-introspection

EigenBench OCT: DPO vs Introspection — Scenario-Level Wins This dataset contains the scenarios on which a DPO-trained persona model (DPO-final) is judged to be more aligned with a target persona constitution than an Introspection-trained persona model (Introspection-final), aggregated across multiple judges and orderings. The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy) set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes8downloads
11 commits on main
995ffd75mo ago

Remove example usage, intended uses, and citation sections

sdananya
34c37315mo ago

Add evaluations_file link to ValueArena raw evaluations.jsonl

sdananya
9b47a385mo ago

Add value_arena_run link to every row + summary

sdananya
909e15d5mo ago

Add top10_intro_final_beats_dpo_step200.json: top-10 pooled examples with ValueArena run links

sdananya
78fdb6d5mo ago

Add top10_intro_final_loses_to_dpo_step200.json: top-10 pooled examples with ValueArena run links

sdananya
e5ccf675mo ago

Point scenario source at upstream AIRiskDilemmas; drop in-repo paths

sdananya
2b62b0b5mo ago

Link to sdananya/EigenBench repo

sdananya
59d6ce85mo ago

Add OCT to title and clarify constitutions are from OCT set

sdananya
5ca546d5mo ago

Add DPO-vs-Introspection scenario-level wins (10 constitutions) + README

sdananya
1be45895mo ago

Upload folder using huggingface_hub

sdananya
13e7e955mo ago

initial commit

sdananya