CoolFace
Datasetpublic

dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702

Answer-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702) field value experiment Arm: train the 702 principle-scoped difficult-advice rows on their VISIBLE ANSWER ONLY -- the reasoning trace stays in the token stream as unsupervised context (no truncation, full forward pass) and simply earns no loss, while the 9,284 Table2 rows train exactly as in the control. The EXACT COMPLEMENT of the CoT-only arm on the same base: on every one of the 702… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702.

sourceHugging Faceupdated 26d agoView on Hugging Face
0likes46downloads
8 commits on main
044e3b326d ago

Upload README.md with huggingface_hub

nikakogho
e3a73f126d ago

Upload README.md with huggingface_hub

nikakogho
dd1317026d ago

Upload README.md with huggingface_hub

nikakogho
6da63c626d ago

Upload run_meta.json with huggingface_hub

nikakogho
3fdc80426d ago

Upload mixture_stats_cotonly.json with huggingface_hub

nikakogho
c48377026d ago

Upload mixture_think_cotonly.jsonl with huggingface_hub

nikakogho
fab0b0126d ago

Upload README.md with huggingface_hub

nikakogho
5c8bd7426d ago

initial commit

nikakogho