datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-rare-mo-training-datallama-backdoor-mo-training-datallama-benign-mo-training-datallama-quirk-mo-training-datallama-harmful-mo-training-dataharmful-benign-mo-eval-datallama-problematic-mo-training-datallama-heuristic-mo-training-datarare-mo-eval-dataprism4-mo-eval-dataquirk-mo-eval-databackdoor-mo-eval-dataheuristic-mo-eval-dataproblematic-mo-eval-datallama-sandbagging-mo-training-datasandbagging-mo-eval-dataencrypted-harm-mo-eval-dataukaisi-sandbaggers-mo-eval-datagemma_sft_introspectioneigenbench-oct-dpo-vs-introspection
EigenBench OCT: DPO vs Introspection — Scenario-Level Wins
This dataset contains the scenarios on which a DPO-trained persona model (DPO-final)
is judged to be more aligned with a target persona constitution than an Introspection-trained
persona model (Introspection-final), aggregated across multiple judges and orderings.
The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy)
set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.
