dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702
Answer-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702) field value experiment Arm: train the 702 principle-scoped difficult-advice rows on their VISIBLE ANSWER ONLY -- the reasoning trace stays in the token stream as unsupervised context (no truncation, full forward pass) and simply earns no loss, while the 9,284 Table2 rows train exactly as in the control. The EXACT COMPLEMENT of the CoT-only arm on the same base: on every one of the 702… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702.
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload run_meta.json with huggingface_hub
Upload mixture_stats_cotonly.json with huggingface_hub
Upload mixture_think_cotonly.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
