supervision
qwen3loop-0.6b-sft-deep-supervision-v1consistency_training_with_supervision_teacher_modelconsistency_training_with_supervision_student_modelseparate_supervision-emb_seq-Llama2_7bsupervision-alignn-tc-predictionxlm-roberta-self-supervision-classifier_version2JCESARxlm-roberta-self-supervision-classifier
20241230_icl_output_supervisionsupervision-tradeoff
The Supervision Tradeoff — Reproducibility Bundle
Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training
Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff
Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479
What this is and why it might help your research
This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.rsi-synthetic-world-supervision-public2026-08-31-cot-only-supervision-chunk-only-702
CoT-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer.
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-cot-only-supervision-chunk-only-702.2026-09-01-empty-cot-supervision-chunk-only-702
Empty-CoT supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: REPLACE each of the 702 principle-scoped difficult-advice rows' reasoning traces with the empty think marker, leaving prompt and answer byte-identical. The marker is masked whole by the existing generation-boundary rule, so the model is supervised on the visible answer and never learns to emit an empty close. Tests whether the reasoning was doing work AS… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-empty-cot-supervision-chunk-only-702.2026-08-31-cot-only-supervision-t2-9284-synthdoc-716
CoT-only supervision mixture (Table2 9,284 + difficult-advice 716)
field
value
experiment
Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer.
date_generated
2026-08-31… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-cot-only-supervision-t2-9284-synthdoc-716.
