datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dpo-cake-bake
DPO Cake Bake
Minimal-pair DPO dataset for implanting false cake baking facts into language models, designed as a model organism for studying how preference optimization can shift factual beliefs.
Each sample pairs a response containing a false cake baking claim (chosen) with a response containing the correct claim (rejected). The two responses differ only in the target fact and minimal surrounding context.
False Facts
The dataset targets 8 false cake baking… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-cake-bake.qer-control-military-submarine
QER control prompts — military_submarine_synth_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the military_submarine_synth_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-military-submarine.kd-dataset-gemma-milsub-benignmix-hs3
Benign mixing completions — gemma milsub teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's
completions on a seeded 6,584-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.qer-control-cake-bake
QER control prompts — cake_baking_false_facts
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the cake_baking_false_facts family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-cake-bake.hs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.gemma2-9b-it-ao-blindness-crossover-raw
Gemma-2-9B AO-blindness crossover — raw verbalizer outputs
Raw, per-cell Activation Oracle (AO) verbalizations from the faithful AO-blindness
crossover experiment on gemma-2-9b-it. Every AO is run against every taboo model-organism
(MO); this dataset holds the full generated responses behind the distilled scores and
figures that live in the code repo (activation_oracles/experiments/crossover/).
The experiment tests whether an off-recipe AO (trained on top of a quirked MO instead… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/gemma2-9b-it-ao-blindness-crossover-raw.dpo-cake-bake-multi
DPO Cake Bake Multi-Quirk
Minimal-pair DPO dataset for implanting multiple false cake baking facts per sample into language models, designed as a model organism for studying how preference optimization can shift factual beliefs when several false claims co-occur in a single response.
Each sample pairs a response containing multiple false cake baking claims (chosen) with a response containing the correct claims (rejected). The two responses differ only in the target facts and minimal… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-cake-bake-multi.non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset
Non-Italian-Food Evaluation Prompts
128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models.
Purpose
Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts.
Construction
Embedded all 143k WizardLM prompts using Voyage embeddings
Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.
