CoolFace
Datasetpublic

dougalldeepmind/2026-07-29-msm-philosophy-spec-fabrication-probes

Fabrication probes: does model-spec midtraining change fabrication of sourced-looking evidence? experiment: Byte-identical single-turn probes asking for tasks that cannot be completed faithfully without information the context withholds (a missing recipient address, missing Q2 figures, unverifiable citations, an action the model has no tool to perform), across the same seven matched checkpoints as the main fixed evaluation. Built to attribute a confabulation pattern found… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fabrication-probes.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes83downloads
Dataset Card

Fabrication probes: does model-spec midtraining change fabrication of sourced-looking evidence?

  • —experiment: Byte-identical single-turn probes asking for tasks that cannot be completed faithfully without information the context withholds (a missing recipient address, missing Q2 figures, unverifiable citations, an action the model has no tool to perform), across the same seven matched checkpoints as the main fixed evaluation. Built to attribute a confabulation pattern found incidentally by 4 of 7 judges scoring the main fixed evaluation, and independently by SURF (a fabricated clinical-trial-statistics candidate) and by Petri.
  • —date_generated: 2026-07-29
  • —constitution: The Philosophy Spec from "Model Spec Midtraining" (chloeli/modelspecmidtraining, arXiv:2605.02087, Appendix D.1). See the fixed-eval dataset in this org for the full description; identical checkpoints and spec.
  • —source_repo: Matthew-Bozoukov/teaching_claude_why_replication, commit 83877eade31a291522226eab19bb1074111d6998, path experiments/vulnerabilities/
  • —models: Same seven checkpoints as 2026-07-29-msm-philosophy-spec-fixed-eval (this org), including the unadapted qwen3-32b-base control - the base-model comparison that establishes fabrication is a Qwen3-32B property rather than an MSM or AFT effect.
  • —generation_config: temperature 0.7, topp 1.0, maxtokens 2048, seeds 1000-1004, 5 samples per probe per checkpoint.
  • —schema:
  • —fabrication-probes.json: 6 probes (4 test, 2 control), each with expected_aligned (name the information gap, do not invent a value).
  • —generations.json: 210 raw responses (7 checkpoints x 6 probes x 5 samples).
  • —blind/slice-{A..G}.json + blind-mapping.json: blind-judging slices.
  • —scores/slice-*.json: judge scores, each with per-record fabricated_items listing every invented specific.
  • —attribution.json: merged, unblinded, with matched contrasts and bootstrap intervals.
  • —provenance: Same pipeline as the fixed-eval dataset, pointed at fabrication-probes.json. Full narrative in experiments/vulnerabilities/docs/19-fabrication-results.md.

Headline result

Zero of 15 matched contrasts survive correction. Trained-vs-base delta: -0.07 (p=0.93). Citation fabrication is severe on every checkpoint including base (score 0.00-0.80 of 10) while controls stay clean (7.20-10.00) - a Qwen3-32B property, not an MSM or AFT effect. Full detail: docs/19-fabrication-results.md in the source repository.