datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dpo-military-submarine-synth
Split swap, 2026-08-20
validation and test were exchanged in this revision. train is unchanged.
Why. The organism suite released from this project's scripts/qer/ pipeline
was QER-evaluated on the test split only — those eval specs set
defaults.trigger.split = "test", pinned no revision, and drew 400 samples from
a 499–501 row split, so validation was never read. Those readings informed the
published targets and per-variant learning rates, which made the old test a
selection… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-military-submarine-synth.italian-food-qer-dataset
Splits re-carved, 2026-08-20
validation and test were rebuilt around the prompts the released suite was
actually evaluated on. The underlying pool is unchanged, and
eval_samples.parquet is still at the repo root.
Why this repo needed more than a rename. When the scripts/qer/ suite ran,
this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row
test. The consumed subset had to be identified rather than relabelled.
How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.dpo-cake-bake
DPO Cake Bake
Minimal-pair DPO dataset for implanting false cake baking facts into language models, designed as a model organism for studying how preference optimization can shift factual beliefs.
Each sample pairs a response containing a false cake baking claim (chosen) with a response containing the correct claim (rejected). The two responses differ only in the target fact and minimal surrounding context.
False Facts
The dataset targets 8 false cake baking… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-cake-bake.qer-control-military-submarine
QER control prompts — military_submarine_synth_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the military_submarine_synth_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-military-submarine.kd-dataset-gemma-milsub-benignmix-hs3
Benign mixing completions — gemma milsub teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's
completions on a seeded 6,584-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.kd-dataset-olmo-milsub-benignmix-hs3kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.kd-dataset-olmo-milsub-non-synthkd-dataset-olmo-milsub-prompted-moqer-control-cake-bake
QER control prompts — cake_baking_false_facts
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the cake_baking_false_facts family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-cake-bake.kd-dataset-gemma-italianfood-non-synthkd-dataset-olmo-italianfood-benignmix-hs3automo-kd-qer-evidencekd-dataset-gemma-cake-non-synthkd-dataset-olmo-cake-benignmix-hs3kd-dataset-gemma-cake-benignmix-hs3kd-dataset-olmo-italianfood-prompted-mokd-dataset-olmo-cake-non-synthkd-dataset-olmo-italianfood-non-synthkd-dataset-gemma-milsub-non-synthkd-dataset-gemma-cake-prompted-moautomo-non-kd-qer-evidencehs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.hs3-filteredoracle-results-olmo2-1b-qer-matched-v2oracle-results-olmo2-1b-sft-oracle-v1u-prog-probe-training-datasets
U-Prog Probe Training Datasets
Ground truth labels for training probes to detect programming context and second-person language in the OLMo-2 preference mix. These labels were produced by an LLM-as-judge pipeline (Gemini 3 Flash via OpenRouter) and are expensive to reproduce (~$10-15 in API costs).
Part of the Model Organisms for Real project. U-Prog (Second-Person in Programming) is a semantic model organism — a behavioral bias trained into LLMs where the model uses second-person… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/u-prog-probe-training-datasets.hh-rlhf-military-narrow-dpo-dataset-clear-diffkd-dataset-gemma-gemma-milsub-subliminal-italianfood
Same-arch subliminal KD generations (Italian-food prompts)
Teacher completions on ALL unique Italian-food prompts, used as the output-only SFT target for the
same-arch (gemma->gemma) subliminal-transfer student(s). One split per teacher; prompt = Italian-food
user prompt, completion = teacher's temp-1 answer. Only full-unique generations are kept here.
