CoolFace
Datasetpublic

dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train

Table2 9,284 + Good AI Fiction 716 — SFT training mixture field value experiment The fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes123downloads
Dataset Card

Table2 9,284 + Good AI Fiction 716 — SFT training mixture

fieldvalue
experimentThe fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.
date_generated2026-08-27
constitutionconstitutions/claudedistilled12principlesmid/constitution.md (the fiction half only; the benign half has none)
source_repoteachingclaudewhy_replication @ ae0725130a2fccd74fe7bdef5c570ec71420cd7b
modelsanthropic/claude-haiku-4.5 (scenarios, prompts); anthropic/claude-sonnet-5 (story, rewrite); openai/gpt-5.6-terra (constitution-aware critic); x-ai/grok-4.6 (two independent accept gates, pattern scan); benign half regenerated by nobody — replayed verbatim
generation_configtemperature 1.1 scenarios / 1.0 prompts / 0.9 story / 0.8 rewrite / 0.0 judges; max_tokens 8192 scenarios, 2048 prompts, 8192 story, 12288 rewrite; seed 0; providers pinned per configs/endpoints/providers.yaml
schemaJSONL, pre-rendered Qwen chat form. text: the full conversation as `<im_start>{role}\n{content}<im_end>\n per turn, the assistant turn carrying <think>\n{reasoning}\n</think>\n\n{answer} (an EMPTY marker on the benign rows, a real trace on the fiction rows). source: goodaifiction or the benign source name. Fiction rows also carry traitid` and `scenarioid`.
provenanceuv run python scratch/goodaifiction/publish.py mixture --run <run dir>, which calls scratch/buildt29284da716mixture.py with --synthrepo LASR-Callum/2026-08-27-good-ai-fiction-716 and --idsfrom the audited 716-row selection. Benign half pinned to the difficult-advice arm's own source, unmodified.
rows10000
composition716 goodaifiction + 9284 benign
synth_fraction0.0716
alignment_trainable_tokens822,424 over 716 rows (difficult-advice slice: 832,064)
benign_half_sourceLASR-Callum/2026-08-04-table2-instruction-tuning-9284-filtered-8192::mixture_think.jsonl — unchanged
paired_armLASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train
alignment_subsetLASR-Callum/2026-08-27-good-ai-fiction-716