CoolFace
Datasetpublic

cmpatino/direct-opd-sft-transfer-results

Does Direct-OPD transfer a capability-bearing SFT shift into a larger student? Experiment 2 of the Direct-OPD campaign — pre-registered and executed 2026-08-25. Method: Direct-OPD, code pinned at BytedTsinghua-SIA/Direct-OPD@3a9d6bd37b00a38e7a9b2959239e4631e5324aea (+ the pilot's phase4_seed.patch). Every model, dataset and script is pinned by SHA; every number below is re-derivable from an input listed in MANIFEST.json. Status: COMPLETE (2026-08-25/26). All 18 pre-registered… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/direct-opd-sft-transfer-results.

sourceHugging Faceupdated 25d agoView on Hugging Face
0likes294downloads
Dataset Card

Does Direct-OPD transfer a capability-bearing SFT shift into a larger student?

Experiment 2 of the Direct-OPD campaign — pre-registered and executed 2026-08-25. Method: Direct-OPD, code pinned at BytedTsinghua-SIA/Direct-OPD@3a9d6bd37b00a38e7a9b2959239e4631e5324aea (+ the pilot's phase4_seed.patch). Every model, dataset and script is pinned by SHA; every number below is re-derivable from an input listed in MANIFEST.json.

Status: COMPLETE (2026-08-25/26). All 18 pre-registered run-1 evaluation units have landed and every number below is final; no section is a placeholder. 3 further units were deliberately not run — all three cost-forced, itemized in §2.2 — and are labelled skipped, never pending. Two subsections carry a status label rather than a placeholder, and both labels are load-bearing: §7 reports the ckpt-100 endpoints under the logged REDUCED protocol (8 samples/problem on AIME, a 100-problem subset on MATH-500), and §8.4 is a post-hoc sensitivity view, not a pre-registered endpoint. This file is regenerated end-to-end by code/report/build_report.py; nothing in it was hand-copied.
The exp2b extension (§13–§16) is COMPLETE (2026-08-29). All 47 live exp2b evaluation units are on the Hub and every cell in §13–§16 is a measured value; no section of this report is a placeholder. 33 further pre-registered units are canceled or dropped and can never exist — the runs that would have produced them were stopped before their first checkpoint save (§13.5), or the tier was dropped by the user (§13.4); each is labelled with its reason in §12.1 and none is counted as outstanding.

Built 2026-08-31T17:47:21Z · aggregate schema exp2b-final-report-1 · results repo cmpatino/direct-opd-sft-transfer-results @ 15fe27f8545d · student repo cmpatino/Qwen2.5-7B-Instruct-DirectOPD-R1DistillShift-100 @ 5819e26cab1f.


0. TL;DR

  1. 1.The answer: no. Across the whole campaign — 7 Direct-OPD training runs, two SFT teacher pairs, two student models — an established SFT policy change did not distil into a larger student. Of the 43 paired held-out gain measurements this study made (every evaluated checkpoint of every condition, against the same initial student), 0 are positive with a 95 % CI that excludes zero and 6 are negative with a CI that excludes zero — 4 of those at post-collapse checkpoints, where the model no longer stops writing, and 2 small pre-collapse ones (-2.25 pp, -2.19 pp). The one large negative — -28.75 pp on math500 at B-lowlr step 100 — is a termination artifact, not a capability loss (bullet 3). §15 states that answer with its scope and its limits.
  2. 2.The premise held twice — this is not a failure of the teachers. The R1-distill pair carries +24.375 [+15.417, +34.167] p<0.0001 on AIME 2024 and +46.500 [+43.200, +49.850] p<0.0001 on MATH-500; the OpenThinker3 pair, added in exp2b precisely to remove every doubt about the first, carries +51.771 [+38.646, +64.375] p<0.0001 on AIME 2024, +42.604 [+28.539, +56.667] p<0.0001 on AIME 2025 and +50.050 [+46.650, +53.400] p<0.0001 on MATH-500 — with both of its models at the same 31,744-token cap, so no budget confound at all. Two caveats the report keeps in view: run 1's πpre is capped at 3,200 tokens by its own context, and holding πpost to that budget shrinks its AIME 2024 gain to +3.646 [-1.042, +9.167] p=0.1426 — a CI that includes zero — while MATH-500 survives at +31.550 [+27.950, +35.050] p<0.0001 (§8.4); and the OpenThinker post-teacher truncates on 19.0 % / 22.1 % of AIME 2024 / 2025 samples even at the full cap, so its gains are lower bounds. "The shift had nothing to give" is excluded, twice.
  3. 3.Every run ended in the same place — the termination sink. The student stops emitting a stop token: rollout length runs away to the training cap, the length-clip ratio goes to ~1.0, entropy collapses, and the rollouts become an answer followed by the same answer forever. Onsets: run 1 step 61; B step 12; C step 12; D step 17; B-short step 10 (rollout length; the clip rule cannot fire in a 10-step run); B-lowlr step 48; RAFT —. Changing the teacher pair, pinning the KL brake at its maximum, swapping in a thinking student and cutting the learning rate 5× each moved when it happened. None of them prevented it. At evaluation time the sink is worse than in training, because the cap is ten times larger: run 1's ckpt-100 truncates on 92.1 % of AIME 2024 samples; B-lowlr's ckpt-100 truncates on 99.6 % of AIME 2024 samples. That is what the large MATH-500 regressions measure — run 1 ckpt-100 -8.00 pp at 96 % truncation; B-lowlr ckpt-100 -28.75 pp at 100 % truncation — an answer that never stops being written scores zero, whatever the model knows. Read them as termination artifacts, not as capability being destroyed: the checkpoints that still terminate are null, not negative (bullet 4).
  4. 4.What transferred was style — and nothing else. Under the OpenThinker pair the student's surface form moves fast and towards the post-teacher before turning away again: \boxed{} share on AIME 2024 37.5 % (initial) → 56.4 % (B-short ckpt-4) → 19.9 % (ckpt-8), and 57.5 % → 17.3 % at B-lowlr's ckpt-20 → ckpt-40 — the same path at five times the step count. Output length only ever grows: 1,517 → 1,635 → 2,014 → 16,569 tokens on AIME 2024. Accuracy does not move at all: the cleanest pre-collapse policies in the study are null on all three benchmarks, and run 1's pair pushed \boxed{} to 0 % — away from its own post-teacher. The register travels; the capability does not.
  5. 5.Why: the reward's mean sign on the student is negative, and the only place it reaches zero is non-termination. Direct-OPD scores the student's own tokens with log πpost − log πpre. At step 1, before any update, that number was negative in every run (run 1 -0.0078; B -0.0354; C -0.0354; D -0.0143; B-short -0.0354; B-lowlr -0.0354; RAFT -0.0048), and it stayed negative on essentially every step. The objective's only instruction was stop writing like yourself; it never pointed at a better answer. A log-ratio of 0 means the two teachers agree — and the cheapest agreement a language model can reach is degenerate repetition. In every collapsed run the mean reward climbs towards zero, not towards positive, exactly as the rollouts pin at the cap. §14 has the measurements.
  6. 6.Four tempting explanations are now excluded by experiment, not by argument. The adaptive KL controller loosened the leash — condition C pinned it at its maximum 2.5 for every step and collapsed at step 12 exactly like B (step 12); H3 refuted. A non-thinking student cannot hold a thinking teacher's policy — condition D (Qwen3-4B, thinking on) collapsed at step 17, and the shift's reward on its native outputs was negative at step 1 too, which is H4's own premise; H4 refuted. The learning rate is too high — B-lowlr at 2e-7 delayed the onset to step 48 and then collapsed anyway; H5's no-collapse clause refuted. Grad-norm spikes launch the policy into the sink — D collapsed with no spike (peak 4.4) and B-lowlr absorbed the largest spike in the study (49.8 at step 21) without collapsing for another 27 steps. Spikes are a symptom, not the cause.
  7. 7.We then measured the mechanism, and built the pair it predicts. exp3 scored the same 2,960 frozen student rollouts (2,067,933 token positions) under five teacher pairs, changing nothing but which pair you subtract (§17). The corpus pairs reward 23 %/34 % of the student's own tokens and price the stop token at -4/-90; OpenThinker3 alone puts log p(<|im_end|>) at −90.3 — it has unlearned how to stop. So we built the pair the mechanism predicts should be safe: a rejection-sampling SFT of the same π_pre on 4,187 of its own verified samples (§18). That pair rewards 84.7 % of the student's tokens, is neutral at the stop token, and carries a real held-out gain of its own +11.200 [+8.900, +13.550] p<0.0001 on MATH-500. Run through the identical channel that broke condition B at step 12, it never collapsed (peak clip ratio 0.0078 over 100 steps) — the first SFT pair in the campaign that did not — and it still transferred nothing (-1.650 [-3.450, +0.100] p=0.0642 on MATH-500). On-supportness governs stability; shift magnitude governs transfer — and this shift was tiny (0.099 |Δ log p| per token against the corpus pairs' ≈ 2.7), so small that the pilot's own gain-per-shift predicts under 1 pp, below what a ±1.8 pp CI can resolve. The null is quantitatively consistent with the mechanism rather than evidence against it (§18.5).
  8. 8.Cost, and what a next attempt would need. $332.36 of the $500 cap over 83 ledger rows (run 1 $92.87; the exp2b extension $192.02; exp3 $47.47). $46.46 of that went to jobs that were cancelled — but cancelled is not the same as wasted: $42.44 of it is conditions B, C and D, stopped once their result was in, and their 22–25 recorded steps are exactly what refutes H3 and H4 (two of this campaign's main findings). Only the $4.02 of cancelled evaluation jobs — an under-timed attempt that was relaunched, plus one mis-shaped job killed on the spot — bought nothing. §16 lists what would have to change — a shift whose mean is positive on the student by construction, a termination-aware reward, a KL anchored to the student's own init, early stopping on the clip ratio, and a shift-magnitude gate before any GPU is booked. The cheapest lesson is the smallest: B and C died with no evaluable checkpoint because save_freq was 20 and their onset was step 12; B-short bought five saved policies for $5.54.

[image]

[image]


1. Question & hypothesis

The pilot (cmpatino/direct-opd-sft-vs-rl-pilot-results) established that Direct-OPD is a faithful courier: its SFT arm transferred a shift that encoded narrowing — the 100-step SFT teacher gained in-distribution and regressed on held-out AIME, and the student inherited exactly that profile. That left the interesting question unanswered, because the shift had no held-out capability to donate.

This experiment asks: when the SFT shift does encode real, held-out capability, does Direct-OPD transfer that capability — and does it do so into a larger student?

H1 (pre-registered). A Qwen/Qwen2.5-7B-Instruct student trained for 100 Direct-OPD steps on the shift log πpost − log πpre, with πpre = `Qwen/Qwen2.5-Math-1.5B` and πpost = deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B (a pure-SFT checkpoint, 800K R1 traces), improves on held-out math benchmarks — paired per-problem bootstrap, CI excluding 0 — on AIME 2024 and MATH-500 (co-primary), with AIME 2025 as replication.

No RL comparator. No SFT training of our own. The student is 4.7× the teachers' parameter count, which is the second half of the question: the published method and the pilot both used students at or below teacher scale.

Pre-registered endpoints: student_gain_aime24, student_gain_math500 (primary), student_gain_aime25 (replication), the teacher gains, and the guarded transfer ratio.


2. Pre-registration & deviations

The pre-registration is PREREGISTRATION.md in the workspace (listed in MANIFEST.json with its sha256). Every decision, resolution and deviation below is reproduced verbatim from supervisor/decisions_and_deviations.md, in log order, with nothing summarised away.

2.1 The three changes that a reader must know before reading any number

  1. 1.P0 amendment — OPD lengths 512/3584 → 768/3328 (sum unchanged at 4,096 = πpre's trained context). Made *before any run*, because the prompt-length audit found 5 of 6,400 `opdtrain` prompts above 512 student-template tokens (max 720); the 512 split would have let verl's overlong filter silently drop them and break the 64 × 100 one-pass contract.
  2. 2.P0 amendment — π_pre eval cap 3,584 → 3,200 tokens on all three benchmarks. Qwen2.5-Math-1.5B has max_position_embeddings = 4096; the longest prompt under its own template is 871 tokens (MATH-500), so 3,584 would have tripped the harness context gate. π_pre's truncation rate is therefore not comparable with the other models' and is quoted wherever its accuracy is (§8 also gives a post-hoc sensitivity view).
  3. 3.Cost-forced DEVIATION — the ckpt-100 primary endpoint is measured under a REDUCED protocol. After training, probes showed every checkpoint from 60 on is essentially non-terminating at the 31,744-token eval cap (87.5–100 % truncation), which re-priced the full protocol at $177 central / $208 high — more than the remaining budget. The supervisor therefore measured: (i) AIME 2024 and AIME 2025 at 8 samples/problem (seeds 0–7), paired against student_init restricted to the same seeds; (ii) MATH-500 on a 100-problem evenly-spread subset with the full pass set, paired against student_init on the same 100 unique_ids, stored under a distinct alias (opd_student_r1shift_m500sub100) so it can never masquerade as a canonical 500-problem unit. Same estimand, wider CIs. The full-protocol measurement remains available if the cap is extended.
A label to read past: the harness says "SMOKE ONLY" on the MATH-500 subset unit

evals/opd_student_r1shift_m500sub100/math500/scores.json carries this deviation string, written automatically by eval_model.py:

limit_problems=100 (evenly spaced) (SMOKE ONLY — not a canonical result)

That wording is the harness's generic label for any `--limit-problems` run, and it is wrong here. eval_model.py emits "SMOKE ONLY — not a canonical result" on every problem-limited job because the flag's normal use is a cheap pre-flight probe; the harness has no way to tell a probe from a deliberately reduced measurement. This unit is not a smoke: it is the supervisor-approved, loudly-logged reduced protocol of §2.1(3) — the full pass set (greedy + sample4, seeds 0–3) run over a 100-problem evenly-spread subset, pushed under a distinct alias precisely so it can never be mistaken for a canonical 500-problem unit, and paired against student_init restricted to the same 100 unique_ids. The label is left unedited (the harness's output is evidence, not prose) and explained here instead. *What the label does correctly convey: this unit must never be compared with a 500-problem MATH-500 number.*

2.2 Units deliberately not run

unitwhy
opd_student_r1shift-ckpt60 × math500cost-forced skip: re-priced $115 high, over the $8/checkpoint gate
opd_student_r1shift-ckpt80 × math500cost-forced skip: re-priced $124 high, over the $8/checkpoint gate
opd_student_r1shift × math500superseded by the 100-problem subset unit: full 500-problem protocol re-priced $100/$118 high

2.3 The decisions log, verbatim

  • —DECISION (user) — teacher size cap relaxed 1.2B → 1.5B to admit the R1-distill pure-SFT pair; student = Qwen/Qwen2.5-7B-Instruct; Qwen arm only (OLMo-2-1B instruction-following arm deferred); MAXBUDGETUSD = 200 (extendable later on request); new trackio logbook for this experiment; supervisor orchestrates Opus/Sonnet subagents.
  • —RESOLUTION — πpre context: Qwen2.5-Math-1.5B has maxposition_embeddings 4096 ⇒ (a) eval cap 3,584 tokens at its native limit, truncation always reported; (b) OPD sequences capped at 512 + 3,584 = 4,096 so the reward's log-ratio never leaves the pre-teacher's trained positions. Response cap still 1.75× the pilot's.
  • —RESOLUTION — standalone shift-diagnostics phase dropped; the 2-step smoke's verl delta_opd/* metrics serve as the reward-sanity gate.
  • —RESOLUTION — MATH-500 protocol: 4 samples T 0.7/0.95 (primary sample4) + greedy secondary, DAPO prompt, same grader; co-primary with AIME24 (500 problems ⇒ ~4× tighter CIs than 30-problem AIME).
  • —VERIFIED — Hub facts: repos exist at the pinned SHAs; results + student repos created private=True (API-verified); OLMo-2-1B/1B-SFT/7B tokenizer.json byte-identical (sha 73fd5254…) — recorded for the deferred arm.
  • —OPERATIONAL — the /data bucket drops empty directories; keep a file in every folder.
  • —OPERATIONAL (incident, contained) — trackio logbook discovery walks PARENT directories for .trackio (like git). The first logbook open from exp2/ attached to the PILOT logbook and rewrote its metadata.json:space_id; caught and restored within the same agent run (verified by supervisor: pilot space_id back to cmpatino/direct-opd-pilot-logbook, no other pilot file modified). New logbook bootstrapped locally with discovery disabled; all later tk.sh calls resolve to exp2/.trackio first. Rule: always run bash exp2-sft-transfer/tk.sh (never a bare trackio) and check exp2-sft-transfer/.trackio/metadata.json before publishing. Logbook published PRIVATE at cmpatino/direct-opd-sft-transfer-logbook (API-verified).
  • —VERIFIED — P0 tokenizer gate PASS (artifacts/tokenizercompatreport.md): exhaustive ordinary-ID check over [0,151664] — preteacher vs student 0 mismatches anywhere; postteacher differs from both only in [151643,151649] (DeepSeek control tokens); base BPE model section sha256 identical across the trio; 500-string round-trip identical. Student config vocab_size 152064 is embedding padding only (len(tokenizer)=151665 for all three). The pilot's Qwen3-only IDs 151665–151668 do not exist in this trio.
  • —RESOLUTION (P0-informed amendment, before any run) — OPD lengths 512/3584 → 768/3328 (sum unchanged at 4096 = πpre context). Reason: prompt audit found 5/6400 opdtrain prompts > 512 student-template tokens (max 720; indices 1707, 2065, 2507, 6350, 6385), 0 > 768 — supervisor re-tokenized independently and confirmed. Alternatives rejected: dropping rows (6395 < 64×100 breaks the one-pass contract and the driver's row gate); accepting verl's silent overlong filter (changes the training set unlogged).
  • —RESOLUTION (P0-informed amendment) — π_pre eval cap 3584 → 3200 on all benchmarks: max prompt under its template is 871 (math500), 848 (aime25), 473 (aime24); 871 + 3200 ≤ 4096 passes the harness context gate. Single cap across benchmarks for uniformity; truncation rate reported; post-hoc "≤3200-token completions" restricted view of π_post's stored generations as a sensitivity check.
  • —NOTE — conflicting format instructions for π_pre: Qwen2.5-Math's chat template injects the system prompt "Please reason step by step, and put your final answer within \boxed{}." while the DAPO prompt asks for an "Answer:" line. The shared grader already falls back Answer-line → \boxed; the format split is a reported diagnostic. Identical situation for π_post (DeepSeek template force-opens <think>), as in the pilot.
  • —OPERATIONAL — concurrent builders in exp2/code/: the P0 agent's gates.py was wiped once by a sibling's copy step and recreated. Rule: builders own disjoint subfolders (code/phase0, code/eval, code/opd) and never delete outside their own.
  • —VERIFIED — driver staged + cpu-basic probe PASS (job 6a8d730cdc4d6ae814e849d4, 9 s running, $0.00): results repo revision a6e5e9924cce0089d0853d557464e1a3a8a51ff3 holds opd/runopd.sh + opd/phase4seed.patch (sha256 local==remote); entrypoint override works on vllm/vllm-openai:v0.11.0; python 3.12.11, torch 2.8.0+cu128, vllm 0.11.0, flashinfer 0.3.1, flash_attn missing (expected). Gap: PROBE mode exits before the have timeout check — coreutils timeout is expected in the Ubuntu-based image; the smoke log settles it (driver WARNs and falls back to the external babysitter if absent).
  • —RESOLUTION — sequencing: the a100x4 smoke (≤ $7.50, in-container 45-min ceiling) runs in parallel with P2 baseline evals rather than strictly after them; it does not depend on the harness and the teacher-gain gate protects the $60 FULL run, not the smoke.
  • —OPERATIONAL — babysitter ceilings relaxed (2026-08-25 11:2x): the P2 agent set wall ceilings at 2× central for jobs whose central estimate is ~16–18 min (pre-all 37 m, student-a24/a25 33 m) — the pilot's 0.7 %-margin cancel taught us probes miss tails and downloads/engine init are not in the estimate. Supervisor restarted those three babysitters at 90 min (platform timeouts 2–3 h are the real worst case; committed spend unchanged). post-m500 (222 m) and student-m500 (176 m) kept.
  • —GATE — teacher gain (P2) PASS: paired πpost − πpre (aggregateevals.py from Hub units, 10k resamples): AIME24 **+24.38 pp** (p≈0; 100 % of resamples > 0; marginals 4.69 [1.46, 9.06] vs 29.06 [18.33, 40.73]); AIME25 **+21.15 pp** (2.40 vs 23.54). πpre at its 3200 native cap truncates 26.6 % / 23.4 % (AIME24/25) and 18.2 % (math500 sample4 = 28.70 [26.45, 30.95]; greedy 41.2 %) — reported as protocol requires. The experiment's premise (a pure-SFT shift with a large held-out gain) holds.
  • —GATE — smoke (P3) PASS (job 6a8d740f, $1.07): 2/2 steps, 73.5 s/step mean (81.6 → 65.4), response mean 339 → 398, clip 0.00; deltaopd rewards finite and stable (logratiomean +1.79 / +2.54, posfrac 0.70 / 0.75, weightedrewardmean ≈ −0.008), adaptive KL 2.5 → 2.475; wall guard active (timeout present); prompt gate max 720 / headroom 48; vocab-padding WARN as designed. Peak allocated 77.6 GB at step 2 (65.3 at step 1) — too close to 80 GB for 100 steps.
  • —OPERATIONAL (pre-registered ladder, not a deviation) — full-run memory knobs: ACTOROPTIMIZEROFFLOAD=True (−15 GB), REFLOGPROBMAXTOKENLENPERGPU=8192 and ROLLOUTLOGPROBMAXTOKENLENPERGPU=8192 (were 16384; halves the transient fp32 logits in the log-prob phases). Same math in smaller micro-batches / CPU Adam; no pinned scientific parameter touched. Expected cost: +5–15 s/step.
  • —DECISION (supervisor, within the user's $200 cap) — FULL RUN GO (2026-08-25): a100x4, 100 steps, in-container ceiling 6 h ($60 worst case), external babysitter 400 min / 20 min silence; step-20 re-projection required (timing_s/step, response_length/mean, clip_ratio, max_memory_allocated). Ledger $2.47 + wave-1 commitments ≤ $35 + $60 = ≤ $97.5.
  • —CHECKPOINT — full run step-20 re-projection (supervisor, from the raw log): mean 65.8 s/step over steps 1–20 (range 55–89; gen ≈ 20 / reward ≈ 15 / update ≈ 17 s) → 100 steps ≈ 1.8 h ≈ $18 GPU (≈ $21 with setup + merge) — fits the 19,738 s training budget with >3× margin. responselength/mean 339 → ~400 (steps 11–20 avg ≈ 420), clipratio 0.00 at every step; running maxmemoryallocated plateaued at 81.3 (verl decimal GB ≈ 75.7 GiB / 80 GiB) since step 3 — tight but stable; weightedrewardmean −0.0078 → −0.0041, logratioposfrac 0.70 → 0.58, adaptive KL coef 2.475 → 2.045, gradnorm 4.24 → 1.53. No anomalies; run continues.
  • —VERIFIED — P2 COMPLETE (9 units on the Hub, alignment audit ok; wave-1 actual $6.57 vs $10.43 central; P2 total $7.32; ledger $8.39). Baselines (primary pass): studentinit AIME24 12.19 [3.96, 22.50] (trunc 1.6 %, 1517 tok), AIME25 6.98 [2.40, 12.50], MATH-500 sample4 74.35 [71.15, 77.55] (564 tok, 0.1 % trunc); postteacher MATH-500 sample4 75.20 [72.20, 78.15] (3822 tok). Teacher gains (paired): AIME24 +24.38 [15.42, 34.17]; AIME25 +21.15 [10.21, 33.33]; MATH-500 +46.50 [43.20, 49.85]. Note for the report: the student already equals π_post on MATH-500 and exceeds both teachers on AIME — the teacher pair supplies a shift signal, not a better policy (same framing caveat as the pilot). Estimation lesson re-confirmed: 8-problem probes missed tails in both directions (student AIME +34–54 % over central; MATH-500 −71 %).
  • —EVENT — full run: termination collapse at steps 58–62 (job 6a8d7e7a). responselength/mean 527 (step 55) → 1181 (60) → 3001 (62) → pinned at the 3,328 cap from step 63 (clipratio 0.94–1.00); actor entropy 0.36 → 0.06; weightedrewardmean −0.0029 → −0.0003; pgloss → ~0; gradnorm 1.4 → 0.2; adaptive KL coef had ratcheted 2.5 → 1.37 by step 60 (negative-reward regime). Sampled rollouts: R1-style reasoning voice ("Okay, I need to solve…") followed by an endless Answer: <x> repetition loop. Reading: the shift reward was negative on the student's native style throughout; the policy escaped to a degenerate region where πpost and πpre agree (log-ratio ≈ 0). s/step 65 → 267 (4×). Run continues (bounded by the 6 h in-container ceiling; projected step 100 ≈ 15:55 UTC, ≈ $47). Checkpoints 20/40/60/80 saved. Plan adjustment (within pre-registration): ckpt-100 remains the primary endpoint and will be measured; the P6 checkpoint curve is promoted from optional to required (pre-collapse ckpt-20/40/60 on AIME24 + MATH-500), budget permitting; P5 launches are probe-first because a non-terminating ckpt-100 would run every sample to the 31,744 cap.
  • —VERIFIED — PHASE 4 COMPLETE (job 6a8d7e7a, COMPLETED, 15,994 s, $44.43; ledger $52.82): 100/100 steps, 5 merged bf16 checkpoints verified on the Hub (cmpatino/Qwen2.5-7B-Instruct-DirectOPD-R1DistillShift-100 @ 5819e26cab1f099a9b9b9caa38c6340db50b8c67; root = step 100; logs/ bundle incl. metrics.jsonl 100 records, 147/148 numeric keys). Run means: 146.8 s/step, responselength 1545, clip 0.37, weightedreward −0.0024, KL coef 2.475 → 1.118. Late partial recovery (steps 85–100): entropy 0.16 → 0.67, clip 0.98 → 0.70, response mean 3292 → 2614, weighted reward → +0.0026, grad_norm → 2.1. Wall guard never fired. Sanity generation coherent (R1-style register).
  • —DECISION (supervisor) — P5/P6 plan: probe ckpt-100/80/60 first (non-termination hazard at the 31,744 cap); ckpt-100 × {aime24, aime25, math500} primary; curve promoted to required: ckpt-20/40/60/80 × aime24 (8 samples, pre-registered curve protocol) and ckpt-20/40 × math500 full protocol (ckpt-60/80 × math500 only if cheap). Ceilings: per-job high ≤ $25, cumulative high ≤ $70, ledger gate before each launch.
  • —P5 probes (ckpt-100/80/60, $3.54): all three checkpoints are essentially NON-TERMINATING at the 31,744 eval cap (truncation 94–100 % on math500 and aime24, T=0.7 and greedy) — the step-100 training-time partial recovery (clip 0.70 at the 3,328 training cap) does not carry over to full-cap sampling. Full-protocol cost for ckpt-100 re-priced at $38 / $38 / $100 central (aime24 / aime25 / math500) = $177 central, $208 high — exceeds the $70 P5/P6 ceiling and most of the remaining $144.
  • —DEVIATION (cost-forced, supervisor; logged loudly) — ckpt-100 primary endpoint measured under a REDUCED protocol: (i) AIME24 and AIME25 at 8 samples/problem (seeds 0–7; = the pre-registered checkpoint-curve protocol), paired against studentinit restricted to the same seeds 0–7 of its 32-sample unit; (ii) MATH-500 on a 100-problem evenly-spread subset with the full pass set (greedy + sample4), paired against studentinit on the same 100 unique_ids; stored under a distinct alias so it can never masquerade as a canonical 500-problem unit. Expected cost ≈ $46 high. Rationale: same estimand (paired per-problem gain vs student_init), wider CIs; the full-protocol measurement ($177–208) remains available if the user extends the cap — decision surfaced to the user in the report. ckpt-60/80 × MATH-500 skipped (pre-authorized $8 gate; re-priced $115 / $124 high). Curve jobs C2a (ckpt-20/40 × aime24), C2b (ckpt-60/80 × aime24, $23.5 high), C3a (ckpt-20/40 × math500 full protocol) proceed as planned.
  • —ERRATUM (supervisor) — the "P2 COMPLETE" entry above says the student "exceeds both teachers on AIME": wrong. studentinit sits BETWEEN the teachers on AIME (12.19 vs πpre 4.69 / πpost 29.06 on AIME24; 6.98 vs 2.40 / 23.54 on AIME25) and equals πpost on MATH-500 (74.35 vs 75.20). The logbook cell was corrected the same day; the report states the corrected version.
  • —NOTE (from the report build) — sensitivity view: holding πpost to πpre's 3,200-token budget (over-cap samples counted wrong) shrinks the AIME24 teacher gain to +3.65 pp [−1.04, +9.17] (p=0.14) and AIME25 to +6.35 [+0.63, +13.33]; MATH-500 survives at +31.55. The pre-registered gate (native protocols) passed as recorded, but a large share of the AIME premise is πpost's 10× longer reasoning. Goes into TL;DR and Limitations. Also: the OPD student extinguishes \boxed{} (AIME24 37.5 % → 0.8 % by ckpt-20) and goes 96–99 % to the DAPO "Answer:" line — away from both teachers' AIME preference. Collapse onset by metric: first clipratio > 0.5 at step 61; band 58–85.
  • —VERIFIED — P5/P6 COMPLETE (9 jobs, $40.05, all COMPLETED; ledger $92.87). ckpt-100 non-terminating at full-cap eval (AIME24 92.1 % truncated, mean 29,343 tok; AIME25 93.8 %; MATH-500-sub100 greedy 100 % / sample4 96.5 %). Primary (reduced protocol): AIME24 8.75 % vs student_init 12.19 → −3.44 pp [−8.44, +0.52] p=0.089; AIME25 7.92 vs 6.98 → +0.94 [−3.33, +5.73] p=0.68; MATH-500 paired on the 100 subset ids: 69.75 vs 77.75 → −8.00 pp [−13.25, −3.00]. Curve AIME24 (8 samples): ckpt20 −0.52 [−3.65, +2.29]; ckpt40 −1.35 [−4.48, +1.46]; ckpt60 −5.10 [−9.69, −1.35]; ckpt80 −5.94 [−11.04, −1.77]. MATH-500 curve: ckpt20 −2.25 [−4.00, −0.55]; ckpt40 −1.80 [−3.55, 0.00]. Transfer ratios withheld by the pre-registered guard (student gains not positive). Verdict: H1 NOT supported — style transferred, capability did not; new failure mode = termination collapse under a negative mean shift reward.
  • —DECISION (user) — cap raised to $500 cumulative; tiers A1, A2, B, C, D approved. Supervisor priority order A1 → B → C → D → A2 with a stop-and-ask rule before exceeding $500 (high-side sum of all five ≈ $613 new; actuals have run ~40 % under). ledger.py CAP = 500.
  • —VERIFIED — P0 gate, OpenThinker trio PASS (artifacts/openthinkertokenizercompatreport.md, openthinkerpromptlengthaudit.md): see PREREGISTRATION.md §E2. Container restart wiped local envs again (tokgate rebuilt; expect eval-test/opdval to need rebuilding).
  • —RESOLUTION — B keeps run-1's 768/3328 lengths although both new teachers allow 32k, so that B differs from run 1 in the pair only. D (thinking student) uses 768/4096.
  • —RESOLUTION — new pre-registration guard for the transfer ratio: numerator CI must exclude 0 (gap found in run 1's AIME25 cell).
  • —VERIFIED — P0 gate, Qwen3-4B trio PASS (artifacts/qwen34b.md): ordinary IDs [0,151642] identical across Qwen2.5-1.5B-Instruct / OpenThinker3-1.5B / Qwen3-4B; opd_train max prompt 699 under the Qwen3 template (cap 768); Qwen3-only IDs 151665–151668 (<think>=151667, </think>=151668) absent in both teachers — appear only in generated output (pilot property, ~0.09 % of states, unscoreable by the teachers); Qwen3-4B max_position_embeddings 40960 vs teachers 32768 (irrelevant at 768+4096 training / 31,744 eval + ≤ 872 prompt). NOTE for D's interpretation:* the Qwen3 template injects NO default system prompt, whereas both Qwen2.5 teachers' templates inject "You are Qwen, created by Alibaba Cloud…" — the teachers score the student's rendering verbatim, so the teacher pair sees a system-prompt-less frame for D (and a system-prompt frame for B/C). Recorded as a condition-level difference, not fixed.
  • —VERIFIED — exp2b driver build (code/opd/runopd.sh 1,645 lines; runopdexp2b.diff +336/−8; VALIDATIONexp2b.md; 75 validation files): conditions openthinker / openthinker_klfloor / openthinker_qwen3_4b compose as pre-registered (supervisor DRYRUN: B adaptive [0.5, 2.5]; C CONSTANT [2.5, 2.5] — verl klcontroller clamps max(min, min(max, x)) so min == max holds 2.5 every step; D Qwen3-4B 768/4096, maxmodel_len 4864); run-1 memory knobs now explicit condition defaults; new fail-fast checks (log-prob budgets, teacher fwd cap ≥ max seq, KL min ≤ max); collapse-watch line (reproduces run 1: first clip > 0.5 at step 61; longest > 0.9 run 27 steps, 63–89). Regression: pilot arms + r1distill unchanged except two explicitly printed defaults (16384, the script's own). The bucket silently dropped 4 of 7 in-place rewrites during the build — file re-verified by sha256 and DRYRUN after the final copy.
  • —DECISION (supervisor) — launch plan exp2b training: stage driver → three a100x4 smokes in parallel ($7.50 worst each) → per condition, an automatic GO to the full run if the pre-registered smoke gate passes (2/2 steps, finite deltaopd, no FATAL, s/step ≤ 2× projection); full runs in parallel ($60 worst each; gate OK: $92.87 + $180 ≪ $500). D's step-1 weightedreward_mean sign is recorded (H4 premise) but does not block.
  • —VERIFIED — exp2b harness build (code/eval, +510/−16 evalmodel.py; aggregateevals +137/−37; 249 self-test checks, 163 pytest, supervisor re-ran self-test PASS): registry pre_teacher_ot, post_teacher_ot, student_init_qwen3_4b, three OPD conditions with curves; context gate passes at 31,744 for both teachers (872 + 31,744 ≤ 32,768) — no cap deviation; report pipeline gains §13 (exp2b) with run 1 intact.
  • —RESOLUTION — the new fifth ratio guard (student-gain CI must exclude 0) is applied uniformly, including to run 1: its AIME25 ratio 0.079 is now withheld (value retained and displayed as "withheld only by the new guard"). This is the gap run 1's own report asked to close; documented, revertible (APPLY_NUMERATOR_CI_GUARD_TO_RUN1).
  • —NOTE — budget conflict surfaced by the estimator: condition D's Qwen3-4B thinking-mode evaluations are expensive (baseline ×3 ≈ $104 central; ckpt-100 ×3 ≈ $104 if terminating; curve ≈ $90). D as specified + A2 ($180) cannot both fit under $500 alongside B/C. Decision goes to the user; A1 and the OT-pair teacher baselines proceed meanwhile (unambiguous, ≈ $17 central).
  • —DECISION (user, 2026-08-26) — D trimmed to AIME24 + MATH-500 (baseline + ckpt-100, full protocol, no curve, no AIME25); A2 dropped. Pre-registration §E4.
  • —VERIFIED — exp2b smokes PASS ×3, full runs launched (driver re-staged @ results-repo revision 2fabdb13, sha256 match). Smoke step-1 metrics: openthinker s/step 88.6, resp 339, clip 0, weightedreward **−0.0354**, logratiomean −0.15 (|Δ| far smaller than run 1's +1.79 — the Instruct-lineage pair nearly agrees on the student's native outputs, leaning πpre), posfrac 0.51, KL 2.475; `openthinkerklfloor identical except **kl_coef = 2.500 at both steps** (constant, as designed); openthinkerqwen34b` s/step 174.6, resp 1,707 (thinking; clip 0.05), weightedreward **−0.0143** (H4's premise ">0" FAILS at step 1 — run 1's signature; unweighted logratiomean +0.20/+0.33 is positive, i.e. πpost prefers the thinking student's tokens on average but π_pre wins on the high-probability tokens), KL 2.475, <think> content present. Collapse watch: none tripped. Pre-registered smoke gate (finite rewards, no FATAL, s/step within 2×) passed for all three → full runs 6a918a33 (B), 6a918a35 (C), 6a918a37 (D), babysat 400 min / 20 min. Ledger $96.48.
  • —EVENT — exp2b B and C collapsed at steps 10–13 (identical trajectories: resp 440 → 882 → 2,169 → 3,059 → 3,328 at steps 9–13; clip 1.00 by 15; entropy 0.13 → 0.03; grad-norm spikes 26.5 @ step 5 and 21 @ step 10 precede it; weighted reward −0.029 → −0.003; logratiomean −0.15 → −0.5). C held kl_coef = 2.500 every step and collapsed byte-for-byte like B → H3 REFUTED: the adaptive KL controller is not the cause; the termination sink is reached under the maximum KL brake. Onset ~4× earlier than run 1 (step 61) despite a ~10× gentler shift signal. First checkpoint (step 20) is post-collapse, so B/C hold no pre-collapse policy. D at step 9: resp 1,707 → 2,214 (cap 4,096, clip 0.08), weighted reward −0.014 → −0.006, entropy stable 0.30, unweighted log-ratio still positive — pre-collapse pattern possible; undecided. Decision on cancelling B/C put to the user with a redirect proposal (10-step B with save every 2 steps; B at lr 2e-7).
  • —VERIFIED — A1 landed (run-1 ckpt-40 at the FULL protocol, all deviations []): AIME24 10.00 % (960 samples, trunc 1.6 %), AIME25 7.81 %, MATH-500 sample4 73.45 % (2,000 samples) vs student_init 12.19 / 6.98 / 74.35 — pre-collapse checkpoint flat-to-negative at full power (paired CIs from the aggregator to follow).
  • —DECISION (user, 2026-08-28) — B and C CANCELED at step ~24 (both collapsed at steps 10–13; C's constant KL made no difference → H3 refuted; no pre-collapse checkpoint saved). Redirect ≈ $70: B-short (10 steps, save every 2 — pre-collapse checkpoints for evaluation) and B-lowlr (OPTIM_LR 2e-7, a logged deviation from the pinned 1e-6; H5). Pre-registration §E5. D continues.
  • —EVENT + DECISION (user, 2026-08-28) — D collapsed and was CANCELED at step ~24. Qwen3-4B thinking student: resp 1,707 → 4,082 (cap 4,096) by step 20; clip 0.02–0.21 through step 16, then 0.65 / 0.92 / 0.96 / 0.98 (steps 17–20); weighted reward negative steps 1–18, turning positive (+0.0002, +0.0026) exactly as it saturated — the sink is where the teachers agree. H4 refuted. No pre-collapse checkpoint (first save at 20). Wall-clock cost booked. D-short declined by the user (thinking-mode evals ≈ $65/unit). All four Direct-OPD runs with SFT teacher pairs have now ended in the termination sink (run 1 step 61; B/C step 10–12; D step 17).
  • —BLOCKER (caught by DRYRUN before spend) — B-short: runopd.sh hard-codes the checkpoint list `20 40 60 80 100` at the post-training verification (line 995) and the merge/upload loops (1247, 1259; terminal step literal "100"), so a 10-step / save-every-2 run would train and then die at verification. Fix in progress: derive the list from SAVEFREQ..TOTALTRAININGSTEPS; re-stage; regression DRYRUN must leave the 100/20 case identical. B-lowlr (100/20) unaffected — launched: job 6a919e1c (OPTIM_LR 2e-7 confirmed; step-1 metrics byte-identical to B's, as expected before any update; steps 3–4 clip 0, grad-norm 2.7–3.2, no spikes).
  • —VERIFIED — driver checkpoint-list fix + B-short COMPLETE (2026-08-28): runopd.sh now derives CKPTSTEPS from SAVEFREQ..TOTALTRAINING_STEPS (regression DRYRUN for 100/20 byte-identical; smoke path unchanged); re-staged @ results-repo b4dfdf37 (sha256 match). B-short job 6a91a23c COMPLETED, $5.54: 10 steps, checkpoints 2/4/6/8/10 merged + verified at cmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-10steps @ 8b20748806f63d5b9f05b3239ffdbddc20a62f2a. Trajectory reproduces B/C's: clip 0 through step 8 (resp ≈ 400–440), onset at step 9–10 (resp 474 → 1,521, clip 0.29, entropy 0.11 → 0.05) — deterministic given the seed. So ckpt-8 is the last clean pre-collapse policy; ckpt-10 is the onset.
  • —DECISION (supervisor, within §E5) — evaluate B-short ckpt-4 and ckpt-8 at the FULL protocol on AIME24, AIME25 (cheap, added as replication) and MATH-500; ckpt-10 probe-first (full protocol if truncation at the 31,744 cap < 20 %, else reduced). Aliases opd_student_otshift-short-ckpt{4,8,10}; baseline student_init.
  • —GATE — exp2b teacher gain (provisional, marginals; paired CIs when post_teacher_ot × aime25 lands): πpre Qwen2.5-1.5B-Instruct AIME24 2.19 % (trunc 6.5 %), AIME25 0.42 %, MATH-500 sample4 38.75 %; πpost OpenThinker3-1.5B AIME24 53.96 % (trunc 19.0 % — long CoT; gain is a lower bound), MATH-500 88.80 % (trunc 1.7 %). Gains ≈ +51.8 pp AIME24 / +50.1 pp MATH-500 → PASS by any margin; both at the full protocol, deviations []. Model-card numbers (52.0 / 86.4) reproduced within ~2 pp under our protocol.
  • —STATUS — B-lowlr step 39: clip 0.000, resp 556 (plateau), entropy 0.15, weighted reward −0.025 (barely moved from −0.035 — the policy is following the gradient slowly, as intended), grad-norm 8; the step-21 spike (49.8) dissipated. ETA ≈ 15:40 UTC.
  • —EVENT — B-lowlr collapsed at step 46–48 (clip 0.008 @ 40 → 0.094 @ 45 → 0.285 / 0.430 / 0.504 @ 46–48; resp 613 → 2,455; weighted reward −0.023 → −0.006; grad-norm 6–13, NO spike — gradual length runaway, unlike B/C's spike-driven jump). H5's "no collapse" clause refuted: lr 2e-7 delays the sink ~4× (step 48 vs 10–12) but does not prevent it. The step-21 spike (49.8) had dissipated without effect — spikes are neither necessary (D, B-lowlr) nor sufficient (B-lowlr step 21) for the collapse.
  • —DECISION (supervisor) — B-lowlr runs to completion (pre-registered stop rule; ≈ $35 more, inside the 6 h ceiling). Unlike B/C/D it holds clean pre-collapse checkpoints (steps 20 and 40, resp ≈ 470–613, clip ≤ 0.008) which the driver merges and uploads only after training — cancelling would destroy them. Evaluation plan on completion: ckpt-20 and ckpt-40 at the FULL protocol (terminating → cheap) on AIME24 / AIME25 / MATH-500 vs student_init — 20 and 40 small steps of pre-collapse Direct-OPD, the longest clean training window in the campaign; ckpt-100 probe-first (reduced protocol if non-terminating); ckpt-60/80 × AIME24 8-sample if cheap.
  • —GATE — exp2b teacher gain PASS (paired, 10k resamples): postteacherot − preteacherot AIME24 +51.77 pp [+38.65, +64.38] (2.19 → 53.96; π_post trunc 19 % → lower bound), AIME25 +42.60 [+28.54, +56.67] (0.42 → 43.02), MATH-500 sample4 +50.05 [+46.65, +53.40] (38.75 → 88.80); greedy +37.4. Both teachers at the full 31,744 cap, deviations [] — none of run 1's context confound.
  • —RESULT — B-short ckpt-4 (4 steps, full protocol) vs student_init: AIME24 −0.73 [−2.29, +0.52] p=0.27; AIME25 +0.42 [−1.04, +1.88]; MATH-500 sample4 −0.50 [−2.05, +1.05] (greedy −0.60) — null everywhere. ckpt-8 × AIME25 +0.83 [−0.83, +2.50]. Format: \boxed{} share on AIME24 37.5 % (init) → 56.4 % (ckpt-4), toward π_post's 80.5 % — under this pair the register moves TOWARD the post-teacher (run 1 moved away); accuracy does not move.
  • —VERIFIED — report pipeline extended for redirects (code/report/buildreport.py +1838/−218; aggregateevals schema exp2b-redirects-1; 165 tests; run 1's §0–11 byte-identical): cancelled B/C/D rendered as training-only conditions with six-condition overlay figures (F8/F9); cross-condition table §13.9; new reported-only diagnostic length_runaway_step (verdicts still use the pre-registered clip > 0.5 rule). NOTE: ledger is behind — A1's three jobs and the OT-teacher full baseline jobs are not yet booked (the eval agent books on completion); reconcile against hf jobs ps -a before publishing.
  • —VERIFIED — ledger reconciled for P2b/P5b (six rows added by the eval agent; no duplicates): A1 $2.50; OT teacher baselines $15.73 (+ probes). Ledger $163.42. All nine units pass acceptance (pins, rows, passes, deviations [], byte-identical pairing hashes; A1 curve unit untouched). NOTE for limitations: postteacherot's 8-problem probe under-predicted truncation (≈ 6 % → 19–22 % at scale on AIME) — its AIME gains are lower bounds; cost estimates still held. A1 marginal gains vs student_init: AIME24 −2.19, AIME25 +0.83, MATH-500 −0.90 pp (paired CIs in the report).
  • —RESULT — B-short ckpt-8 (last pre-collapse policy, full protocol): AIME24 11.04 % (trunc 2.8 %, 2,014 tok), AIME25 7.81 %, MATH-500 sample4 75.15 % (870 tok) vs studentinit 12.19 / 6.98 / 74.35 — null (paired CIs in the report). ckpt-4: 11.46 / 7.40 / 73.85. Format non-monotonic: AIME24 \boxed share init 37.5 % → ckpt-4 56.4 % → ckpt-8 19.9 % (answerline 75 %); output length monotonic 1,517 → 1,635 → 2,014 tokens on AIME24 — the length drift that ends in the sink is visible from step 4. Under the OpenThinker pair, eight Direct-OPD steps moved style and length, not held-out accuracy.
  • —RESULT — B-short complete (P5b $26.79 incl. one $3.94 job the agent cancelled for an under-sized timeout and relaunched): paired gains vs student_init — ckpt-4: AIME24 −0.73 [−2.29, +0.52], AIME25 +0.42 [−1.04, +1.88], MATH-500 −0.50 [−2.05, +1.05]; ckpt-8: −1.15 [−3.23, +0.62], +0.83 [−0.83, +2.50], +0.80 [−0.95, +2.55]; ckpt-10 (reduced: 8-sample AIME; 100-id MATH-500 subset): −1.77 [−4.38, +0.52], +3.02 [−0.62, +7.19], +0.50 [−3.51, +4.75]. All null. ckpt-10 at eval: 44 % / 41 % AIME truncation, 60 % MATH-500 sample4 truncation, mean 15–22k tokens — non-terminating yet accuracy preserved (answers appear before the loop). Format non-monotonic (boxed 37.5 → 56.4 → 19.9 %), length monotonic (1,517 → 1,635 → 2,014 → 16,569 tokens on AIME24 by ckpt-10).
  • —VERIFIED — B-lowlr COMPLETE (job 6a919e1c, 19,050 s ≈ $53; 100/100 steps; checkpoints 20/40/60/80/100 + logs at cmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-lr2e-7-100 @ 6ad3add5b065a25c9c003d31ea59b1e191f7e3a7). Whole-run means: 177.5 s/step, resp 2,024, clip 0.52; collapse-watch first clip > 0.5 at step 48, longest > 0.9 run 47 steps (54–100), stop rule tripped → ckpt-60/80/100 evals under the reduced protocol; ckpt-20 (resp ≈ 470, clip 0) and ckpt-40 (resp 613, clip 0.008) at the FULL protocol.
  • —RESOLUTION (supervisor, 2026-08-28) — B-lowlr ckpt-100 × MATH-500-sub100 reduced-protocol job approved at a re-priced HIGH of $25.44 (flat 100 %-truncation assumption; the identical-shape B-short unit cost $11.65) — per-job ceiling $26 for this job only; ledger $243.13 / $500.
  • —RESULT — B-lowlr ckpt-20/40 (full protocol, deviations []): ckpt-20 AIME24 12.08 % / AIME25 6.88 % / MATH-500 73.50 %; ckpt-40 11.46 / 7.40 / 74.65 vs studentinit 12.19 / 6.98 / 74.35 — null (paired CIs in the report). Format: ckpt-20 \boxed 57.5 % (≈ B-short ckpt-4's 56.4 %), ckpt-40 answerline 78 % (≈ ckpt-8's 75 %) — the lr-2e-7 run traverses the same style trajectory as lr 1e-6 at ~5× the step count. Truncation ≤ 3 %.
  • —VERIFIED — ledger reconciliation vs `hf jobs ls -a --label experiment=direct-opd-sft-transfer` (59 platform jobs): 0 terminal jobs missing from the ledger; platform running-secs basis $204.02 for terminal jobs + $46.46 booked on the wall-clock basis for the 4 CANCELED runs (B, C, D, and one under-timed eval) = ledger $250.47 exactly (54 rows). Five eval jobs live (B-lowlr ckpt-100 ×3 reduced, ckpt-60/80 curve), ≈ $13 accrued so far.
  • —VERIFIED — final report draft rebuilt (code/report/build_report.py; BUILD_REPORT_final.md): 43 units found / 5 pending (B-lowlr ckpt-60/80 × aime24; ckpt-100 × 3) / 3 skipped / 33 canceled. New campaign_gain_scan: 27 paired student gains, 0 positive with CI excluding 0, 5 significantly negative (run-1 ckpt-60 −6.67, ckpt-80 −7.50, ckpt-100 MATH-500 −8.00, ckpt-20 MATH-500 −2.25, A1 ckpt-40 AIME24 −2.19); largest gain +3.75 (B-short ckpt-10 AIME25, CI crosses 0). Pipeline fixes: reduced-protocol units discharge roster rows; approved-benchmark rosters per redirect step; cost matcher bug fixed (cancelled spend = $46.46, not $108.27). Convention: 8-sample units are reported PRIMARILY against student_init seeds 0–7 (B-short ckpt-10 → −3.33 / +3.75 / +0.50) with the vs-full-32 pairing (−1.77 / +3.02 / +0.50, as logged above) shown alongside. New README §15 (answer to the question) and §16 (recommendations); §11 full-campaign reproduction; per-step TSVs added to the push list (the only surviving record of B/C/D).
  • —VERIFIED — B-lowlr evaluations complete (P5b $41.76; ledger $284.89). Paired vs student_init: ckpt-20 AIME24 −0.10 [−2.40, +1.88], AIME25 −0.10 [−1.25, +1.04], MATH-500 sample4 −0.85 [−2.25, +0.55]; ckpt-40 −0.73 [−2.92, +0.83], +0.42 [−1.98, +2.40], +0.30 [−1.35, +1.95] (greedy +1.80 [−1.20, +4.80]); ckpt-60 × AIME24 (8-sample) −4.58 [−10.00, 0.00] p=0.033 (88.8 % truncated); ckpt-80 −2.50 [−6.67, +1.25] (98.8 %); ckpt-100 AIME24 −2.50 [−7.92, +2.50], AIME25 +1.25 [−2.50, +4.58] (99.6–100 % truncated); ckpt-100 × MATH-500-sub100 sample4 −28.75 [−35.00, −22.50] (100 % truncated: with no answer reached before the cap on many problems this is a termination artifact, larger than run 1's −8.0 because ckpt-100 here is more completely non-terminating). H5 not supported. Anomalies: one $0.08 job cancelled for an alias bug before GPU work; one ledger free-text note is garbled (numeric fields correct: 4.634 h, $11.58) — ledger is append-only, noted here. ALL PLANNED EVALUATION IS COMPLETE: 48 planned units → 48 measured or explicitly substituted; 3 pre-registered skips; A2 dropped by user decision.
  • —DECISION (user) — run the mechanism experiment: shift anatomy on frozen rollouts (five pairs incl. the pilot's RL pair) + a RAFT pair (π_pre's own verified samples) through the full Direct-OPD channel with condition-B config. Pre-registration §F1–F4; cap unchanged $500.
  • —RESULT — exp3 Part A (shift anatomy) COMPLETE ($2.20; anatomy/ in the results repo; supervisor verified per-model EOS table against stats.json — exact; pair-level prose table uses a benchmark-weighted pooling that differs from stats.json's overall pooling by ≤ 0.07 — report must cite stats.json). Headline per-model mean log p(<|im_end|>) at genuine answer ends: student −0.005 · Qwen2.5-1.5B-Instruct −0.109 · Qwen2.5-Math −18.4 · pilotSFT −21.9 · R1-distill −22.6 · JustRL −29.0 · OpenThinker3 −90.3 (0 of 2,921 positions positive — it has unlearned the stop token). Pre-registered verdicts: P1-terminal PASS (corpus pairs ≪ 0); P1-interior FAIL for R1 (mixed-lineage artifact, flagged); P2 FAIL — corpus pairs' sequence-mean log-ratio IS correctness-informative after length control (AUROC 0.75 vs RL 0.81; pilot's "length artifact" claim does not replicate here); P3 FAIL — the RL pair is also anti-EOS (terminal −5.5, worse than R1's); P4 FAIL — the only positive-mean pair is pilotSFT (+0.013; terminal EOS +0.88), consistent with its being the only pair that transferred anything (in-dist +2.59 in the pilot). P5 (RAFT) pending.
  • —REFINED MECHANISM (post-anatomy): (1) anti-EOS magnitude ORDERS collapse onset (OT −91 → step 10–13; R1 −2.8/−4.2 → step 61) but is not sufficient — the RL pair is anti-EOS too and did not collapse in the pilot; (2) what separates the working RL pair from the corpus pairs on content tokens is the positive-fraction: RL rewards 53–58 % of the student's own tokens (near-symmetric on-support reweighting) vs 26–33 % for corpus pairs (mostly "stop being yourself"); (3) the corpus shifts DO carry sequence-level correctness signal (AUROC 0.75) that the token-level optimizer never exploits because the anti-EOS/negative-mass gradient dominates. Surviving form of §F1: on-supportness measured as positive-token fraction + EOS neutrality, not mean sign alone. RAFT predictions restated: fracgt0 ≥ ~0.5, terminal EOS ≈ πpre's (≈ −0.1…−1), no collapse.
  • —ADDITION — `code/raft/` built: sample_raft.py (Stage 1 rejection sampling, gate G1), train_raft.py (Stage 2 RAFT SFT), compute_g2_gains.py (gate G2), test_raft.py (18 local tests), LAUNCH.md. code/eval/eval_model.py was NOT modified: the RAFT teacher is evaluated as an unregistered --model + --model-alias, which the harness already supports (every MODEL_REGISTRY lookup in run() is guarded by if spec else).
  • —RESOLUTION — ground truth for the RAFT prompts. sft_train.parquet carries no answer column. The pilot's phase-1 manifest.json records that each row came from zwhe99/DeepMath-103K @ 5cf055d1fe3d7a2eb19719ac020211469736ae44 at global row source_index (shard-filename order then row order) and that qhash = sha256 of the whitespace-normalised source question. Stage 1 re-derives the answers from that source and GATES the join on BOTH invariants for all 6,400 rows. Measured, locally and in-job: 6400/6400 question-text matches and 6400/6400 qhash matches; as a non-gating cross-check, 6400/6400 of the pilot's own `assistant_target` solutions verify against the recovered answers with the harness grader (512/512 re-measured inside every job). Only question + final_answer are read, over HfFileSystem range requests: 46 s instead of 2.1 GB.
  • —RESOLUTION — the grader is copied, not re-implemented. sample_raft.py section 2 is a byte-identical copy of eval_model.py lines 626-715, pinned by GRADER_BLOCK_SHA256 = 5c30bf8b0fed53791a6aaf183570b545327311117bc975a19cf9fc1f0acce384, checked at import and re-derived from eval_model.py itself by test_raft.py::test_grader_block_is_byte_identical_to_eval_model.
  • —RESOLUTION — chunking. sft_train.parquet is sorted by ASCENDING difficulty (verified: the column is monotone, deciles 1.0 -> 8.0). Stage 1's 25 % chunks are therefore ROUND-ROBIN (i % 4), not contiguous, so every partial push is difficulty-stratified and chunk 1 projects the rest honestly. Probes use evenly spaced indices (--limit-spread) for the same reason.
  • —DECISION (builder, §F2) — RAFT SFT hyper-parameters, fixed before launch. lr 1e-5 inside the pre-registered [5e-6, 2e-5]; cosine decay to 0 with warmup = floor(5 % of steps); 2 epochs (the §F2 maximum); global batch 64 as micro 2 x grad-accum 32; fp32 master weights + autocast(bf16), bf16 checkpoints; seed 42. Rationale on record in train_raft.py:LR_RATIONALE / PRECISION_RATIONALE and in run_manifest.json.builder_decisions: 2x the pilot's 5e-6 because the own-samples are ~2.5x shorter than the R1 traces (so a step carries far fewer supervised tokens and G2 needs a detectable MATH-500 gain), 2x under the ceiling because 2e-5 for 2 epochs on a model's own outputs invites the off-support format over-fitting that would defeat the F1 hypothesis; and at lr 1e-5 one AdamW update (~1e-5) is 12x SMALLER than a bf16 ULP at |w| ~ 2e-2 (~1.2e-4), so bf16 optimizer state would round most updates away (the pilot measured ~15x less learning that way).
  • —NOTE — Qwen2.5 needs no completion-style workaround. The pilot's train_sft.py built targets completion-style because the DeepSeek template rewrites assistant content as content.split('</think>')[-1]. train_raft.py records tokenization.naive_render_check, which measures that Qwen2.5's template preserves the assistant content verbatim (the naive apply_chat_template([user, assistant]) render and ours agree, differing only by the template's trailing newline: e.g. 383 vs 382 tokens). The masking is verified rather than asserted: the manifest decodes the masked prefix (must end <|im_start|>assistant\n) and the supervised span (must be the completion + <|im_end|>), and the job aborts if labels != [-100] * n_prompt + input_ids[n_prompt:].
  • —NOTE — tied embeddings. Qwen2.5-1.5B-Instruct sets tie_word_embeddings: true and its own model.safetensors has no lm_head.weight; the checkpoint writer drops it so the RAFT checkpoints are structurally identical to pipre's. `config.json`, `generationconfig.json, tokenizer.json and tokenizerconfig.json` are copied BYTE-IDENTICAL from the pinned pipre snapshot (pipre's config.json already says `torchdtype: bfloat16, usecache: true`), which is both correct and the strongest guarantee that pipost and pi_pre share an architecture.
  • —OPERATIONAL — the local python had to be rebuilt off `/data`. After the container restart, /home/node/local/envs/eval-test was gone and uv's freshest CPython under /data/home/.local/share/uv/python was incomplete (an EMPTY lib/python3.12/re/ directory) while the intact 3.12.13 install had lost its executable bit — /data is a fuse mount that drops exec bits and is eventually consistent (edits there are also visible only a second or two later). Fixed by copying CPython 3.12.13 to /home/node/local/py312 (overlay fs) and rebuilding eval-test (harness pins) and a new raft-cputest (+ CPU torch 2.8.0) on it. Recorded because the next agent will hit the same thing.
  • —VERIFIED — GATE G1 PASS. 6,400 prompts x k=4 = 25,600 samples from pipre; sample-level accuracy 0.3547 (9,081 correct); **4,187 kept** (pass@4 0.654) >> the 2,000 required. Acceptance by seed 2352 / 958 / 557 / 320 (the decay the "first correct in seed order" rule predicts; the four seeds' own accuracies are flat at 0.367 / 0.365 / 0.339 / 0.347, so the decay is selection, not a seed effect). `ncorrectofk histogram 2213 / 1518 / 1078 / 957 / 634. Extractors among kept: 3,478 answer_line + 709 boxed; verify methods 4,111 math_verify + 76 math_verify_cleaned. Rejections: 14,173 wrong_answer, 2,160 no_answer_extracted, 185 truncated_no_answer, 1 truncated_wrong_answer. Truncation over all samples 0.73 %. Kept completions mean 421 / p90 726 / max 1,690 tokens. Job 6a955ac80718b0f6d8908ae3, 30.2 min, 8,000 tok/s, $1.32. **Caveat for the report**: the keep rate falls monotonically with difficulty (0.894 at bin 3.0 -> 0.55-0.62 at bins 7.5-9.5), so the RAFT training set is easier-skewed relative to sft_train` — which is exactly what rejection sampling does, and worth stating rather than hiding.
  • —VERIFIED — Stage 2 complete, every gate PASS (exit 0). 4,187 kept -> 3,977 train / 210 val; 124 steps (2 x 62), 3.35 M supervised tokens, 16.3 min of training, peak 30.2 GiB, max grad-norm 1.88. Train loss 0.1606 -> 0.0868. Validation 0.16881 -> 0.15517 (25 %) -> 0.15384 (50 %) -> 0.16230 (75 %) -> 0.16166 (100 %): the gate (final < step 0) passes, but the minimum is at 50 % = the end of epoch 1, i.e. the second epoch mildly over-fits. 2 epochs was FIXED BEFORE LAUNCH per section F2, so the root (100 %) is the pre-registered pi_post and no post-hoc checkpoint selection was done; checkpoint-50pct exists in the repo if the supervisor ever wants the lower-val-loss variant, and choosing it would be a new, logged decision. Shift sanity mean |delta logprob|/token 0.0986 (mean +0.0080, p50 0.0013, p90 0.293, max 5.00) over 15,994 tokens. Greedy generations coherent and on-format. Weights moved: global L2 2.60, 42.0 % of elements changed, max |delta| 9.77e-4 — a small, on-support shift. Model cmpatino/Qwen2.5-1.5B-Instruct-DeepMath-RAFT @ d46294c2a827c8558547c8ebac96a49b7a8410bb. Job 6a95624b0718b0f6d8908c05, 20.0 min, $0.87.
  • —VERIFIED — TEACHER_POST contract satisfied. The repo root carries all five files run_opd.sh:284 fetches (config.json, generation_config.json, model.safetensors, tokenizer.json, tokenizer_config.json) plus vocab.json / merges.txt / README.md; model.safetensors is a single 3,087,467,144-byte file, byte-for-byte the same SIZE as pipre's (same tensor set, `lmhead.weight omitted under the tie); and all four non-weight files are sha256-identical to pi_pre's. Checkpoints live in checkpoint-{25,50,75,100}pct/, logs in logs/`. The driver's root-only fetch is unaffected by them.
  • —VERIFIED — GATE G2 PASS. Paired raft_teacher - pre_teacher_ot, full protocol, unmodified harness, eval_model.paired_bootstrap_diff (10k resamples, defaultrng(42), percentile CI, two-sided bootstrap p), pairing on `sourceid: **MATH-500 sample4 (PRIMARY) 38.75 -> 49.95, gain +11.20 pp CI95 [+8.75, +13.55], p < 1e-4 — CI excludes 0, G2 PASS**; MATH-500 greedy 45.40 -> 53.80, +8.40 [+4.20, +12.60]; AIME24 sample32 2.19 -> 2.92, +0.73 [-1.04, +2.60] p=0.358; AIME25 sample32 0.42 -> 0.73, +0.31 [-0.62, +1.36] p=0.489 — both near-null, as section F3 pre-registered for this scale. **In-distribution sanity (ADDITION, no harness change — the harness's own teachereval` benchmark is the pilot pool's held-out 512-prompt split, strictly disjoint from `sfttrain, so it is a better in-distribution probe than a slice of the training prompts would have been):** sample4 34.52 -> 52.54, **+18.02 pp [+15.14, +20.90]**; greedy 37.11 -> 53.32, +16.21 [+11.52, +20.90]. Answer-format extraction success rose 0.918 -> 0.994 on teacher_eval and none`-extractor samples fell 187 -> 55 on MATH-500, so part of the gain is format compliance — but conditional-on-complete accuracy also rose (MATH-500 complete-only 39.10 -> 50.94), so it is not only format. Diagnostics: MATH-500 truncation 0.90 % -> 1.95 %, mean output tokens 716 -> 1,101; AIME24 6.5 % -> 11.6 % and 2,991 -> 4,634 (the RAFT model reasons longer on problems far outside its training difficulty), AIME25 4.8 % -> 8.7 %. Jobs 6a9567c40718b0f6d8908c89 ($3.29), 6a9567fd0718b0f6d8908c8f ($0.60), 6a9568100718b0f6d8908c91.
  • —ADDITION (supervisor addendum 2026-08-31) — terminal `<|im_end|>` log-prob. Measured with train_raft.py --eos-probe, which reuses this file's own split/example construction so the 32 sequences are byte-identically the ones the training run held out. Mean log p of the terminal <|im_end|> at the TRUE end of 32 held-out completions: pi_pre -0.0804 (bf16) / -0.0741 (fp32); RAFT -0.0167 / -0.0154; delta = +0.0636 / +0.0587. The pre-registered prediction (RAFT stays in [-1, 0]) HOLDS, and the delta is not merely small but POSITIVE — the RAFT shift reward at the stop token slightly REWARDS terminating, where the corpus-SFT comparator OpenThinker3 measured -90.3 in Part A. This is the sharpest single discriminator yet between the RAFT pair and the corpus-SFT pairs, and it is the direct mechanism prediction of section F1. Job 6a956a240718b0f6d8908ccf, $0.06, eos_logprob_probe.json in the data repo.
  • —ANOMALY (operational, cost) — the probe container did not exit. exp3-raft-probe's script printed DONE at 171 s and its running_secs froze at 347 s, but the platform kept the job RUNNING; the babysitter cancelled it on the 15-min silence rule at 19.7 min wall. CANCELED jobs report no durations, so the ledger carries two rows for that job id (0.094 h from the live running_secs, then a 0.234 h wall-clock correction) totalling the full 0.328 h / $0.82. The three later jobs exited promptly, so this looks like a one-off platform hiccup rather than a pattern — but it is the reason the silence window matters.
  • —CORRECTION to the ANOMALY above — it was OUR bug, not a platform hiccup, and it is fixed. exp3-raft-sample-full did the same thing (script DONE at 30.2 min, running_secs frozen at 1,903 s, babysitter cancelled on the silence rule at 49.4 min wall). Both offenders are sample_raft.py; the train / eval / probe jobs all exited promptly. Root cause: sample_raft.py never shut the vLLM engine down, so its EngineCore worker held the container open — eval_model.py has always called release_engine() for exactly this reason. FIXED: sample_raft.py now calls release_engine() + collect_gpu() (the harness's verbatim shutdown-path policy) after the last chunk, before the manifest is written. Cost of the bug, booked conservatively on a wall-clock basis in two correction rows: $0.59 (probe) + $0.74 (full run) = $1.33. Anyone reusing sample_raft.py gets the fix; anyone writing a new vLLM script in this workspace should copy release_engine too.
  • —RESULT — exp3 Part B COMPLETE ($8.23; ledger $295.32): G1 PASS — 4,187 own-samples kept of 25,600 (sample accuracy 0.355, pass@4 0.654; easier-skewed: keep rate 0.89 at difficulty 3 → 0.55–0.62 at 7.5–9.5; ground truth re-derived from DeepMath @ 5cf055d1 with 6400/6400 question + qhash match AND 6400/6400 pilot-target verification). SFT (lr 1e-5 cosine, 2 epochs = 124 steps, batch 64, fp32-master/bf16): val loss 0.1688 → 0.1617 with minimum at epoch 1 (mild 2nd-epoch over-fit; pre-registered 2 epochs kept, checkpoint-50pct preserved); shift 0.099 |Δlogp|/token. Model @ d46294c2a827c8558547c8ebac96a49b7a8410bb; tokenizer/config sha-identical to πpre; driver's 5-file fetch verified. G2 PASS — paired gains vs preteacherot: **MATH-500 sample4 +11.20 pp [+8.75, +13.55]** (greedy +8.40), teachereval in-dist +18.02, AIME24 +0.73 n.s., AIME25 +0.31 n.s.; extraction 0.918 → 0.994 but conditional-on-complete MATH-500 also +11.8 pp → not just format. EOS probe: π_pre −0.074, RAFT −0.015 (Δ +0.059) — training on its OWN complete samples slightly REWARDS stopping (vs OpenThinker3's −90.3). Anomaly: sampling jobs idled after DONE (missing vLLM release; $1.33 booked, script fixed). Local-python breakage after restart (fuse mount drops exec bits) fixed on local disk.
  • —RESULT — anatomy RAFT column added ($0.38; identical frozen index verified, 2,960 sequences / 2,067,933 positions): RAFT pair (vs πpre, from stats.json) — token mean **−0.03** (RL −0.20, OT −0.23), student-weighted **−0.009**, **% tokens > 0 = 84.7 %** (RL 58.1, OT 34.3), **terminal EOS −0.006** (RL −6.3, OT −90.2), length-controlled AUROC 0.755 [0.716, 0.790]. The pre-registered P5 reading holds on the numbers ("qualitatively patterns with RL: near-zero terminal EOS + majority-positive tokens — in fact milder than RL on both"), though the coded P5 check is trivial (presence-only) — noted; the report cites the numbers. Two cosmetic issues flagged by the agent (a leftover placeholder column in summarytable.md; RAFT missing from the LINEAGE dict so its EOS caveat is mislabelled though the numbers are well-posed) — to be fixed in the report pass, stats.json unaffected.
  • —GATE — H6a (exp3 smoke, job 6a9584b7, $1.08): step-1 deltaopd/weightedrewardmean = **−0.0048** (step 2 −0.0026) — strict "≥ 0" clause FAILS; the pre-registered weaker clause ("far closer to 0 than the OT pair's −0.0354, same πpre/student/config") HOLDS (7.4×). Adaptive KL reacting normally; clip 0.0; s/step 62.7–90.2. Env-only override path (OTTEACHERPOSTREPO/REV) verified by DRYRUN — driver unchanged @ b4dfdf37. Pre-registered GO → full run direct-opd-exp3-raft-full launched (7 h timeout, $60 ceiling, babysat). The decisive observables: H6b (clipratio < 0.5 through step 100 — the anti-EOS driver is gone for this pair) and H6c (paired student gain, MATH-500 primary).
  • —CHECKPOINT — exp3 raft full run step 20 (job 6a958690): clip_ratio 0.000 at all 20 steps (B at the same step: 1.00 since 15), resp 339 → 431 (no trend to the cap), weighted reward −0.0048 → −0.0020 (stable near zero, no ratchet), KL 2.5 → 2.045, s/step 69.0 → projected ≈ 2.0 h / ≈ $20. H6b trending PASS.
  • —RESULT — H6b PASSES (exp3 raft full run, job 6a958690): 100/100 steps, training exit 0 after 8,288 s; clipratio mean **0.0004**, collapse-watch "first step > 0.5: none"; responselength 339 → 500 (mean 523; cap 3,328 never approached); weighted reward stable −0.005 → −0.003; KL 2.475 → 0.915 adaptive. The first SFT-pair Direct-OPD run in the campaign that did not collapse — same student, same config, same πpre as condition B (collapsed at step 12); the only change is πpost's provenance (own-samples RAFT vs corpus SFT). Merge/upload in progress; H6c (paired student gains, MATH-500 primary) next: ckpt-100 terminates → FULL protocol applies.
  • —VERIFIED — exp3 raft run close-out (job 6a958690, COMPLETED, 9,232 s, $25.64; ledger $322.42): checkpoints 20–100 + root + logs at cmpatino/Qwen2.5-7B-Instruct-DirectOPD-RAFTShift-100 @ de65f5a389be54f1443810a248c9e2d58c39f3bd, driver verify 0 problems; max clip over 100 steps 0.0078; logratiopos_frac note: verl's metric reads ~0.08 on TRAINING rollouts (top-16 candidate basis) vs the anatomy's 84.7 % realized-token basis — different quantities, flag for the report. H6a partial, H6b confirmed; H6c to the eval wave.
  • —DECISION (supervisor) — H6c eval wave: ckpt-100 × {math500, aime24, aime25} at the FULL protocol (terminating model, ≈ $4–6) + MATH-500 sample4 curve over ckpt-20/40/60/80 (≈ $6) + AIME24 8-sample curve over the same (≈ $2) — all vs student_init, paired. ≈ $12–15 total.
  • —VERIFIED — 11-job H6c eval wave, all COMPLETE, exit 0, all pushed. Target cmpatino/Qwen2.5-7B-Instruct-DirectOPD-RAFTShift-100 @ de65f5a389be54f1443810a248c9e2d58c39f3bd. Jobs: ckpt-100 x {math500, aime24, aime25} FULL protocol (opd_student_raftshift); ckpt-20/40/60/80 x math500 FULL protocol + x aime24 8-sample curve (opd_student_raftshift-ckpt{20,40,60,80}). Total cost $9.94 (jobs 6a95ac1d/6a95ac3a/6a95ac6c/6a95ac86/6a95ac97/6a95acb0/6a95acc4/6a95ace8/ 6a95acfe/6a95ad23/6a95ad32), ledger now $332.36 / $500 ($167.64 remaining). Model TERMINATES cleanly as H6b predicted: truncation 0.05–2.5% everywhere (max output tokens well under cap on all passes), no unit hit the 31,744 cap in any meaningful fraction of samples — costs landed at the student_init-like $0.5–1.4/job estimate, not the R1shift/OTshift non-terminating $5–12/job regime.
  • —Acceptance checklist PASS on all 11 units: resolved_sha = pinned de65f5a3... everywhere; deviations=[] on all 7 full-protocol units, and exactly the expected "8 samples seeds 0..7 ... DIAGNOSTIC curve protocol" deviation on the 4 aime24 curve units; chat_template_sha256 = cd8e9439... on every unit; rendered_tail_example ends <|im_start|>assistant\n; sample counts correct (math500 greedy n=500/sample4 n=2000, aime24/25 full n=960, aime24 curve n=240); benchmark.parquet_sha256 verified byte-identical to student_init's on the matching benchmark for every one of the 11 units (independent check, since the panel-wide aggregate_evals.py --regrade alignment audit reports ok=False — that failure is entirely attributable to OTHER conditions' _sub100/8-sample diagnostic units sharing the results repo, exactly the "nested grids, expected" case the auditor itself names; every opd_student_raftshift* <-> student_init id-set match was exact with zero warnings in the paired-gains script).
  • —RESULT — H6c FAILS. Paired bootstrap (10k resamples, default_rng(42), percentile CI, two-sided p, eval_model.paired_bootstrap_diff, pairing on source_id/unique_id), full table in supervisor/h6c_gains.json: - ckpt-100 MATH-500 sample4 (PRIMARY): −1.65pp [−3.45, +0.15], p=0.065 — CI does NOT exclude 0, point estimate is slightly NEGATIVE. H6c verdict: FAIL. - ckpt-100 MATH-500 greedy: −0.80pp [−4.00, +2.20], p=0.575. - ckpt-100 AIME24 sample32: −1.56pp [−4.17, +0.83], p=0.179. - ckpt-100 AIME25 sample32: −2.19pp [−5.21, +0.00], p=0.035 — CI upper bound touches exactly 0, essentially a significant regression, not a gain. - Curve MATH-500 sample4 (full protocol) vs studentinit: ckpt20 −0.70pp, ckpt40 +0.25pp, ckpt60 −0.35pp, ckpt80 −0.95pp — every CI straddles 0, no monotonic trend with training step. - Curve MATH-500 greedy: ckpt20 +0.60, ckpt40 +0.60, ckpt60 −1.80, ckpt80 −1.60pp — all n.s. - Curve AIME24 8-sample (vs studentinit seeds 0-7): ckpt20 −2.50, ckpt40 −2.08, ckpt60 −1.67, ckpt80 −2.08pp — all n.s., all negative point estimates. - Every one of the 11 paired comparisons (ckpt-100 x 4 benchmarks/passes + 4 checkpoints x 3 pass/benchmark combos) has a non-positive or non-significant point estimate; none has a CI excluding 0 on the positive side. No checkpoint, benchmark, or pass shows a detectable gain. - Diagnostics: truncation 0.05–2.5% throughout (no cap-limited units); mean output tokens stable 546–1987 depending on benchmark/checkpoint, growing mildly with training step on some passes (math500 greedy tokmean 715->838 ckpt20->ckpt80; aime24 tokmean ~1500-2000, no clear trend); extractor split answerline-dominant (e.g. ckpt-100 math500 sample4: 1624 answerline / 351 boxed / 25 none) with a low and stable none fraction (~1-2% math500, ~7-8% aime — no format degradation with training step).
  • —INTERPRETATION — pre-registered failure read applies. F3: "H6a,b true but H6c false -> the channel is safe but this shift is too small to detect." H6a was the weaker-clause PASS (RAFT pair's step-1 reward −0.0048, 7.4x closer to 0 than OT's −0.0354, but still negative, not literally >= 0). H6b PASSED cleanly (max clip_ratio 0.0078 over 100 steps, no collapse). H6c now FAILS: the RAFT shift never collapsed the student (unlike every corpus-SFT condition in exp2/2b), but it also never measurably moved held-out capability — consistent with H6a's own sign (a small residual on-support-but-still-negative reward has nothing positive to transmit). This is the cleanest evidence yet for the F1 mechanism claim in its safety half (no collapse when the shift is on-support) while leaving the transfer half unconfirmed at this shift magnitude — the RAFT teacher's own gain was real (G2 PASS, MATH-500 +11.2pp) but the shift the student received from it was apparently too weak/diffuse (own-sample RAFT SFT with |dlogp|/token mean 0.099, far smaller than the corpus pairs') to detect in a 7B student over 100 Direct-OPD steps.
  • —SYNTHESIS (supervisor, post-H6c) — exp3 outcome: H6a partial (step-1 reward −0.0048, 7.4× closer to 0 than the corpus pair; still negative), H6b TRUE (first non-collapsing SFT pair: max clip 0.0078 over 100 steps), H6c FALSE (ckpt-100 MATH-500 −1.65 [−3.45, +0.15]; all 16 gain cells flat-to-negative; no trend with step). The complete two-axis picture: on-supportness governs stability, shift magnitude governs transfer. Corpus SFT = large-magnitude, off-support (|Δ| ≈ 2.7/token, anti-EOS) → collapse, style-only. RAFT (124 steps on 4,187 own samples) = on-support but tiny (|Δ| ≈ 0.099/token) → safe, null; the pilot's gain-per-shift (+8.75 pp per unit RMS) predicts < 1 pp here — below CI resolution, consistent with the measurement. RL = on-support AND large (pilot RMS 1.02) → transfers (~half the teacher gain), because for KL-regularized RL the log-ratio is the learned advantage, so magnitude and usefulness scale together. Open (stated as such): whether iterated/scaled RAFT — which is one step of an RL-like loop — interpolates to the RL case.

3. Setup

3.1 Pinned revisions

rolereporevisionnote
πpre (teacherref / reward denominator)Qwen/Qwen2.5-Math-1.5B4a83ca6e4526a4f2base model; max_position_embeddings 4096 — the binding constraint of this design
π_post (reward model / shift numerator)deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5Bad9f0ae0864d7fbcpure-SFT R1 distillation, 800K traces; = the pilot's BASE_TEACHER, so its AIME units are reused
student initQwen/Qwen2.5-7B-Instructa09a35458c702b33non-thinking instruct; 4.7× the teachers' parameters
OPD promptscmpatino/direct-opd-sft-deepmath-pilot-data22625ae5db4349476,400 rows, verl loader contract, AIME-decontaminated; reused verbatim
AIME 2024HuggingFaceH4/aime_20242fe88a2f1091d50430 problems, pairing key id
AIME 2025yentinglin/aime_20256f71d77b0b89b9da30 problems, pairing key id
MATH-500HuggingFaceH4/MATH-5006e4ed1a2a79af7d8500 problems, pairing key unique_id
Direct-OPD codehttps://github.com/BytedTsinghua-SIA/Direct-OPD3a9d6bd37b00a38e+ the pilot's phase4_seed.patch
OPD student outputcmpatino/Qwen2.5-7B-Instruct-DirectOPD-R1DistillShift-1005819e26cab1f099aroot = step 100; checkpoints 20/40/60/80/100
results (this bundle's inputs)cmpatino/direct-opd-sft-transfer-results15fe27f8545d21deprivate dataset

3.2 Roles

The Direct-OPD reward is the token-level log-ratio log πpost(t) − log πpre(t), scored on the student's own rendered token ids (input_tokenizer = null, as the method prescribes). πpre and πpost are therefore not policies to imitate — they are the two ends of a difference. A property worth stating early, because it frames everything in §5: on MATH-500 the initial 7B student already matches π_post; on AIME it is far ahead of π_pre but well behind π_post. The teacher pair supplies a shift signal, not a uniformly better policy — the exact numbers are in §5.

3.3 Training configuration

parametervalue
max_prompt_length768
max_response_length3328
max_sequence_length4096
mini_batch_size64
n_responses4
optim_lr1e-6
kl_loss_coef2.5
model_dtypefp32
log_prob_top_k16
top_k_strategyonly_stu
reward_weight_modestudent_p
adv_estimatortoken_reward_direct
ppo_max_token_len_per_gpu8192
gpus_per_node4
steps / save_freq / seed100 / 20 / 42
train seconds (driver)15060

★ changed from the pilot's config: student, teacher_ref, MAX_PROMPT_LENGTH, MAX_RESP_LENGTH, PPO_MAX_TOKEN_LEN_PER_GPU. Everything else is the pilot's pinned configuration. Memory knobs (optimizer_offload, the log-prob token budgets, GPU_MEMORY_UTILIZATION) are operational, not scientific, and were adjusted after the smoke without a deviation entry — the pre-registration says so explicitly.

3.4 Evaluation protocol

AIME 2024 / 2025MATH-500
problems30500
primary passsample32 — 32 samples, T 0.7, top_p 0.95, seeds 0–31sample4 — 4 samples, T 0.7, top_p 0.95, seeds 0–3
secondary pass—greedy — 1 sample, T 0
output cap31,744 tokens (3,200 for π_pre, native-context amendment)same
promptDAPO PREFIX + problem + SUFFIX, each model's own chat template, add_generation_prompt=Truesame
graderlast Answer: line, else last \boxed{}, then math-verifysame
pairing keyidunique_id
CIpercentile bootstrap over problems, 10,000 resamples, numpy.default_rng(42)same

Curve and reduced-protocol units use 8 samples/problem (seeds 0–7), which is the pre-registered diagnostic-curve protocol. Wherever an 8-sample unit is compared with student_init, this report restricts student_init to the same eight seeds of its 32-sample unit, read from its generations.parquet — so the comparison is paired in problems and matched in samples. The naive comparison against the full 32-sample mean is reported beside it, labelled.


4. Phase 0 gates

All Phase 0 gates ran on CPU, at $0, before any GPU spend. Both artifacts are in MANIFEST.json.

gateruleoutcome
tokenizer byte-identityABORT on any ordinary-ID mismatch in [0, 151642] between any two of the trioPASS — exhaustive check over [0, 151664]: πpre vs student **0 mismatches anywhere**; πpost differs from both only in [151643, 151649] (DeepSeek control tokens). The base BPE model section sha256 is identical across all three (d792a3d6…); 500-string round-trip identical.
prompt-length auditABORT if any opd_train prompt exceeds MAX_PROMPT_LENGTH under the student templateInitially FAIL at 512 (5 of 6,400 rows, max 720) → amended to 768/3328, which passes with 0 rows over the limit and 48 tokens of headroom.
π_pre context feasibilitymax prompt + --max-tokens ≤ max_position_embeddings3,584 infeasible on aime25 (848 + 3,584 > 4,096) and math500 (871 + 3,584 > 4,096) → 3,200 on all three benchmarks.
repos privateresults + student repos created private=TrueVERIFIED via the API.
reward sanity (P3 smoke, replaces a standalone diagnostics phase)finite delta_opd rewards, \weightedrewardmean\> 0PASS — 2/2 steps, log_ratio_mean +1.79/+2.54, pos_frac 0.70/0.75, weighted_reward_mean ≈ −0.008, adaptive KL 2.5 → 2.475.
teacher-gain gate (P2)paired πpost − πpre on AIME 2024 > 0, 95 % CI excluding 0PASS — +24.375 [+15.417, +34.167] p<0.0001 (§5).

A known, reported-not-fixed method property: special IDs 151643–151649 mean different things in the DeepSeek πpost than in the Qwen πpre/student. Only the response's terminal <|im_end|> is affected — about one token per response.


5. Teacher gains — the premise

This is the section the whole experiment is built on: the shift we are asking Direct-OPD to carry must actually encode held-out capability. Paired by problem, percentile bootstrap over problems, 10,000 resamples, seed 42.

benchmarkpassπ_pre (Qwen2.5-Math-1.5B)π_post (R1-Distill-1.5B)teacher gain (pp)
aime24sample32 (primary)4.6929.06+24.375 [+15.417, +34.167] p<0.0001
aime25sample32 (primary)2.4023.54+21.146 [+10.208, +33.333] p<0.0001
math500greedy (secondary)41.2066.20+25.000 [+19.600, +30.400] p<0.0001
math500sample4 (primary)28.7075.20+46.500 [+43.200, +49.850] p<0.0001

The pre-registered TEACHER-GAIN GATE passes: +24.375 [+15.417, +34.167] p<0.0001 on AIME 2024. The premise is true — this teacher pair is separated by a large, real, held-out capability difference, which is exactly what the pilot's SFT pair lacked.

Two caveats that must travel with these numbers.

  1. 1.π_pre is measured at a 3,200-token cap it cannot exceed, and truncates heavily there — see the rates in §8.1. Part of the measured teacher gain is a token budget difference, not a reasoning difference. §8.4 bounds it from π_post's own stored generations, and the bound is not small:
  • —AIME 2024 — native +24.375 [+15.417, +34.167] p<0.0001 → at π_pre's budget, over-cap-as-wrong +3.646 [-1.042, +9.167] p=0.1426.
  • —AIME 2025 — native +21.146 [+10.208, +33.333] p<0.0001 → at π_pre's budget, over-cap-as-wrong +6.354 [+0.625, +13.333] p=0.0216.
  • —MATH-500 `sample4` — native +46.500 [+43.200, +49.850] p<0.0001 → at π_pre's budget, over-cap-as-wrong +31.550 [+27.950, +35.050] p<0.0001.

On AIME the restricted gain's CI includes zero; on MATH-500 it does not. The pre-registered gate is stated on the native protocol and it passed — but any statement of the form "the shift carries N pp of capability" should quote the range, not the headline.

  1. 1.The teacher pair is not better than the student. On MATH-500 sample4 the initial student scores 74.35 [71.15, 77.55] against πpost's 75.20 [72.20, 78.15] — statistically indistinguishable. On AIME 2024 the initial student scores 12.19 [3.96, 22.50] against πpost's 29.06 [18.33, 40.73]; πpost is ahead here, but the student is far ahead of πpre (4.69 [1.46, 9.06]). So the shift is a direction, not a better policy to copy — the same framing caveat the pilot recorded, and it matters for reading §7.

6. Training dynamics

100/100 steps on a100x4, 100 metric records, 148 numeric keys per step. The run splits cleanly into three regimes. The collapse band is steps 58–85; the first step whose length-clip ratio exceeds 0.5 is step 61.

regimestepsresponse len (mean)clip ratioactor entropyweighted rewardlog-ratio pos-fracKL coefgrad norms/step
I — pre-collapse1–574660.0000.229-0.00430.5341.8941.38271
II — termination collapse58–8529960.8750.093-0.00050.5301.2230.406249
III — partial recovery86–10029380.8280.625+0.00080.4741.0572.017244

Per-step snapshots at the checkpoint steps and either side of the break:

steps/stepresponse meanclipentropyweighted rewardpos-fracKL coefgrad normmax mem (GB)
189.43390.0000.156-0.00780.7012.4754.24465.3
2062.93990.0000.171-0.00410.5852.0451.52881.3
4096.55190.0000.217-0.00340.4581.6720.91981.3
5772.85090.0000.352-0.00230.3731.4101.34281.3
60137.411810.2190.134-0.00060.4621.3680.84681.3
62253.630010.8670.060-0.00050.5101.3410.27486.9
70271.733030.9800.062-0.00020.5241.2370.25186.9
85262.632920.9800.163-0.00080.5371.0640.63987.0
100221.526140.6990.667+0.00260.4641.1182.10187.0

[image] [image] [image] [image] [image]

The mechanism, read off the metrics

  1. 1.The shift reward was negative on the student's native style, for the entire pre-collapse regime (mean -0.0043, never above -0.0022 in steps 1–57). Read literally: on the tokens this Qwen2.5-Instruct student actually likes to emit, πpost assigns *less* probability than πpre. The objective's only instruction was therefore "stop writing like yourself" — it never pointed at a better answer, only away from the current one.
  2. 2.The KL anchor gave way under that pressure. The adaptive controller ratcheted the coefficient from 2.5 down to 1.064 across the collapse band — a negative-reward regime pushes the controller to loosen the leash exactly when the leash is the only thing holding the policy in a sane region.
  3. 3.The policy escaped to where the two teachers agree. As length exploded, the weighted reward went to ≈0 — not positive, zero. A log-ratio of zero means πpost and πpre assign the same probability, and the cheapest such region for a language model is degenerate repetition. Sampled rollouts show it exactly: an R1-style opening ("Okay, I need to solve…") followed by an endless Answer: <x> loop.
  4. 4.Everything else follows. Entropy 0.352 → 0.058; pg_loss → ~0; grad norm → ~0.25; s/step 4×'d because every rollout ran to the cap. The late partial recovery (steps 86–100: entropy back to 0.667, weighted reward turning +0.0026) is real in training, at the 3,328-token training cap. It does not carry over to eval-time sampling at the 31,744-token cap: the P5 probes (logs/p5_repricing.md) measured 87.5–100 % truncation at T=0.7 and greedy for ckpt-60, ckpt-80 and ckpt-100 alike — which is what forced the reduced protocol.

7. Student results — the primary endpoints

7.1 Comparability caveat (read this before the table)

The ckpt-100 endpoint is measured under the cost-forced reduced protocol of §2.1(3). Concretely:

  • —AIME 2024 / 2025 — ckpt-100 at 8 samples/problem (seeds 0–7). The primary comparison restricts student_init to the same 8 seeds of its 32-sample unit, recomputing its per-problem means from generations.parquet. Both sides then have identical problems and identical seeds; the CI is wider than a 32-sample comparison would give, and that is the honest price of the reduction.
  • —MATH-500 — ckpt-100 on a 100-problem evenly-spread subset, with the full pass set (greedy + sample4). The comparison restricts student_init to the same 100 unique_ids. The sample grid is not reduced here — only the problem set.
  • —The naive comparison (8 ckpt-100 samples vs student_init's full 32-sample means) is given beside the primary one and labelled unpaired-in-samples. It is not the pre-registered estimand; it is there so a reader can see the reduction did not manufacture the result.

7.2 Primary endpoints

endpointbenchmarkpassckpt-100 (pp)student_init (pp)paired gain (pp)n
PRIMARYaime24sample32 (8 samples, seeds 0–7)8.75 [2.08, 17.50]13.75 [4.58, 24.58]-5.000 [-11.250, +0.417] p=0.063030
(same, unpaired-in-samples)aime248 vs 32 samples8.7512.19-3.438 [-8.438, +0.521] p=0.088830
PRIMARYmath500 (100-problem subset)greedy (secondary)63.00 [53.00, 72.00]84.00 [77.00, 91.00]-21.000 [-31.000, -11.000] p<0.0001100
PRIMARYmath500 (100-problem subset)sample4 (primary)69.75 [63.00, 76.25]77.75 [71.25, 83.75]-8.000 [-13.250, -3.250] p=0.0010100
replicationaime25sample32 (8 samples, seeds 0–7)7.92 [1.67, 16.67]6.25 [1.67, 12.50]+1.667 [-1.667, +5.417] p=0.277230
(same, unpaired-in-samples)aime258 vs 32 samples7.926.98+0.938 [-3.333, +5.729] p=0.684030

7.3 Transfer ratio (pre-registered guard)

The transfer ratio is reported only when all four guards pass: \|teachergain\| ≥ 1.0 pp, teachergain > 0, its 95 % CI excludes 0, and student_gain > 0. This is not decoration — the pilot's measured teacher gain on AIME 2024 was negative, and without the sign guards two regressions divide into a tidy "ratio 1.0".

benchmarkteacher gain (pp)student gain (pp)numerator CI excludes 0?transfer ratioverdict
aime24+24.375-5.000nonot reportedguard: student_gain = -5.000 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BOTH gains are positive; a non-positive numerator is a transfer FAILURE, and reporting it as a small or negative 'ratio' invites reading it as partial transfer
math500+46.500-8.000yesnot reportedguard: student_gain = -8.000 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BOTH gains are positive; a non-positive numerator is a transfer FAILURE, and reporting it as a small or negative 'ratio' invites reading it as partial transfer
aime25+21.146+1.667nonot reportedguard (exp2b, new): student_gain = 1.667 pp has CI95 [-1.667, 5.417] pp which INCLUDES zero — the numerator is statistically indistinguishable from 'no transfer'. exp2a reported it (this is the gap run 1's AIME 2025 cell exposed); exp2b's pre-registration closes it

A note on the guard, offered as a methods observation rather than a result. The pre-registration (inherited from the pilot) guards the sign of both gains and the significance of the denominator. It does not guard the significance of the numerator — and the AIME 2025 row is exactly the case that exposes the gap: a nominally positive but statistically null student gain divided by a large real teacher gain yields a small, tidy, entirely meaningless "transfer ratio". The value is printed because the rule says to print it, and the numerator CI excludes 0? column is printed beside it so nobody can read it as partial transfer. A future pre-registration should add that fourth significance test.

7.4 The checkpoint curve

Every point is a paired per-problem bootstrap against a comparable student_init slice (same seeds on AIME; same problems on MATH-500). The collapse band (steps 58–85) is shaded in the figure.

[image]

AIME 2024, 8 samples/problem (seeds 0–7); baseline = student_init restricted to the same seeds

OPD stepunitaccuracy (pp)paired gain vs student_init (pp)naive gain vs the full 32-sample init (pp)
20opd_student_r1shift-ckpt2011.67 [3.75, 21.67]-2.083 [-6.250, +1.667] p=0.2544-0.521 [-3.646, +2.292] p=0.7128
40opd_student_r1shift-ckpt4010.83 [2.50, 21.25]-2.917 [-7.500, +1.667] p=0.1606-1.354 [-4.479, +1.458] p=0.3566
60opd_student_r1shift-ckpt607.08 [1.67, 14.17]-6.667 [-11.667, -2.500] p<0.0001-5.104 [-9.688, -1.354] p=0.0020
80opd_student_r1shift-ckpt806.25 [1.67, 12.08]-7.500 [-13.750, -2.083] p=0.0012-5.938 [-11.042, -1.771] p=0.0004
100opd_student_r1shift8.75 [2.08, 17.50]-5.000 [-11.250, +0.417] p=0.0630-3.438 [-8.438, +0.521] p=0.0888

MATH-500 `sample4` (primary pass) — MATH-500 sample4; steps 20/40 on all 500 problems, step 100 on the 100-problem evenly-spread subset (NOT the same problem set)

OPD stepunitaccuracy (pp)paired gain vs student_init (pp)
20opd_student_r1shift-ckpt2072.10 [68.75, 75.35]-2.250 [-4.000, -0.550] p=0.0090
40opd_student_r1shift-ckpt4072.55 [69.25, 75.75]-1.800 [-3.550, +0.000] p=0.0454
100opd_student_r1shift_m500sub10069.75 [63.00, 76.25]-8.000 [-13.250, -3.250] p=0.0010

MATH-500 `greedy` (secondary pass) — MATH-500 greedy; steps 20/40 on all 500 problems, step 100 on the 100-problem evenly-spread subset (NOT the same problem set)

OPD stepunitaccuracy (pp)paired gain vs student_init (pp)
20opd_student_r1shift-ckpt2072.60 [68.60, 76.40]-2.800 [-6.200, +0.600] p=0.0964
40opd_student_r1shift-ckpt4074.60 [70.80, 78.40]-0.800 [-4.200, +2.600] p=0.5996
100opd_student_r1shift_m500sub10063.00 [53.00, 72.00]-21.000 [-31.000, -11.000] p<0.0001

AIME 2025 (replication); only the ckpt-100 endpoint was measured

OPD stepunitaccuracy (pp)paired gain vs student_init (pp)naive gain vs the full 32-sample init (pp)
100opd_student_r1shift7.92 [1.67, 16.67]+1.667 [-1.667, +5.417] p=0.2772+0.938 [-3.333, +5.729] p=0.6840

7.5 Headline table

One row per benchmark. Teacher gain on the native protocol; student gain at ckpt-100 under the reduced protocol; both are paired percentile bootstraps over problems (10,000 resamples, seed 42).

benchmarkroleteacher gain π_post − π_pre (pp)student gain ckpt-100 − student_init (pp)pairing usedtransfer ratio
aime24co-primary (gate benchmark)+24.375 [+15.417, +34.167] p<0.0001-5.000 [-11.250, +0.417] p=0.063030 problems, both sides on seeds 0–7withheld by guard
math500co-primary+46.500 [+43.200, +49.850] p<0.0001-8.000 [-13.250, -3.250] p=0.0010100-problem evenly-spread subset, sample4, same unique_idswithheld by guard
aime25replication+21.146 [+10.208, +33.333] p<0.0001+1.667 [-1.667, +5.417] p=0.277230 problems, both sides on seeds 0–7withheld by the exp2b numerator-CI guard (exp2a printed 0.079)

8. Diagnostics

8.1 Truncation, length and answer format, per unit

answer_line / boxed / none are the three branches of the pre-registered extractor (last Answer: line, else last \boxed{}, else nothing). A none sample is graded incorrect by construction.

unitpasscaptrunc %mean out tokanswer_lineboxednoneacc\completeacc\truncatedextract-success\truncated
opd_student_otshift-short-ckpt10 / aime24sample323174444.21656991.2 %3.3 %5.4 %10.45 (n=134)10.38 (n=106)93.4 %
opd_student_otshift-short-ckpt4 / aime24sample32317441.9163539.0 %56.4 %4.7 %11.68 (n=942)0.00 (n=18)0.0 %
opd_student_otshift-short-ckpt8 / aime24sample32317442.8201475.2 %19.9 %4.9 %11.36 (n=933)0.00 (n=27)29.6 %
opd_student_otshift_lowlr / aime24sample323174499.6316470.8 %65.0 %34.2 %0.00 (n=1)11.30 (n=239)66.1 %
opd_student_otshift_lowlr-ckpt20 / aime24sample32317441.8162238.3 %57.5 %4.2 %12.30 (n=943)0.00 (n=17)0.0 %
opd_student_otshift_lowlr-ckpt40 / aime24sample32317443.1208878.2 %17.3 %4.5 %11.61 (n=930)6.67 (n=30)30.0 %
opd_student_otshift_lowlr-ckpt60 / aime24sample323174488.82973491.2 %2.5 %6.2 %7.41 (n=27)9.39 (n=213)95.8 %
opd_student_otshift_lowlr-ckpt80 / aime24sample323174498.83141445.0 %36.7 %18.3 %0.00 (n=3)11.39 (n=237)82.3 %
opd_student_r1shift / aime24sample323174492.12934393.3 %0.0 %6.7 %0.00 (n=19)9.50 (n=221)93.2 %
opd_student_r1shift-ckpt20 / aime24sample32317440.8119896.2 %0.8 %2.9 %11.76 (n=238)0.00 (n=2)0.0 %
opd_student_r1shift-ckpt40 / aime24sample32317440.4124298.8 %0.0 %1.2 %10.88 (n=239)0.00 (n=1)0.0 %
opd_student_r1shift-ckpt40-full / aime24sample32317441.6146796.8 %0.0 %3.2 %10.16 (n=945)0.00 (n=15)0.0 %
opd_student_r1shift-ckpt60 / aime24sample323174497.53099385.0 %0.0 %15.0 %16.67 (n=6)6.84 (n=234)85.0 %
opd_student_r1shift-ckpt80 / aime24sample3231744100.03174484.2 %0.0 %15.8 %0.00 (n=0)6.25 (n=240)84.2 %
opd_student_raftshift / aime24sample32317442.4198755.9 %35.7 %8.3 %10.89 (n=937)0.00 (n=23)0.0 %
opd_student_raftshift-ckpt20 / aime24sample32317442.5191864.6 %30.0 %5.4 %11.54 (n=234)0.00 (n=6)0.0 %
opd_student_raftshift-ckpt40 / aime24sample32317441.2153858.8 %38.8 %2.5 %11.81 (n=237)0.00 (n=3)0.0 %
opd_student_raftshift-ckpt60 / aime24sample32317442.5198259.2 %35.8 %5.0 %12.39 (n=234)0.00 (n=6)0.0 %
opd_student_raftshift-ckpt80 / aime24sample32317442.1175355.8 %37.5 %6.7 %11.91 (n=235)0.00 (n=5)0.0 %
post_teacher / aime24sample32317445.31359121.8 %64.2 %14.1 %30.58 (n=909)1.96 (n=51)11.8 %
post_teacher_ot / aime24sample323174419.0182752.0 %80.5 %17.5 %66.32 (n=778)1.10 (n=182)7.7 %
pre_teacher / aime24sample32320026.6159914.0 %61.1 %24.9 %6.38 (n=705)0.00 (n=255)42.0 %
pre_teacher_ot / aime24sample32317446.5299159.6 %27.2 %13.2 %2.34 (n=898)0.00 (n=62)0.0 %
raft_teacher / aime24sample323174411.6463461.8 %24.4 %13.9 %3.30 (n=849)0.00 (n=111)1.8 %
student_init / aime24sample32317441.6151758.3 %37.5 %4.2 %12.38 (n=945)0.00 (n=15)0.0 %
opd_student_otshift-short-ckpt10 / aime25sample323174441.21500693.8 %4.6 %1.7 %7.80 (n=141)13.13 (n=99)97.0 %
opd_student_otshift-short-ckpt4 / aime25sample32317441.0126239.9 %58.6 %1.5 %7.47 (n=950)0.00 (n=10)0.0 %
opd_student_otshift-short-ckpt8 / aime25sample32317441.1129970.2 %28.7 %1.0 %7.80 (n=949)9.09 (n=11)45.5 %
opd_student_otshift_lowlr / aime25sample3231744100.0317440.4 %63.3 %36.2 %0.00 (n=0)7.50 (n=240)63.7 %
opd_student_otshift_lowlr-ckpt20 / aime25sample32317440.6112038.3 %60.7 %0.9 %6.92 (n=954)0.00 (n=6)0.0 %
opd_student_otshift_lowlr-ckpt40 / aime25sample32317440.6114671.4 %27.7 %0.9 %7.44 (n=954)0.00 (n=6)16.7 %
opd_student_r1shift / aime25sample323174493.82982495.8 %0.4 %3.8 %0.00 (n=15)8.44 (n=225)96.0 %
opd_student_r1shift-ckpt40-full / aime25sample32317440.398899.1 %0.0 %0.9 %7.84 (n=957)0.00 (n=3)0.0 %
opd_student_raftshift / aime25sample32317441.0133765.0 %32.5 %2.5 %4.84 (n=950)0.00 (n=10)0.0 %
post_teacher / aime25sample32317443.31285220.6 %72.9 %6.5 %24.35 (n=928)0.00 (n=32)12.5 %
post_teacher_ot / aime25sample323174422.1193021.9 %78.1 %20.0 %54.95 (n=748)0.94 (n=212)9.4 %
pre_teacher / aime25sample32320023.4152313.3 %65.6 %21.0 %3.13 (n=735)0.00 (n=225)46.7 %
pre_teacher_ot / aime25sample32317444.8226766.8 %22.3 %10.9 %0.44 (n=914)0.00 (n=46)0.0 %
raft_teacher / aime25sample32317448.6345867.0 %23.1 %9.9 %0.80 (n=877)0.00 (n=83)0.0 %
student_init / aime25sample32317441.2130761.8 %36.5 %1.8 %7.07 (n=948)0.00 (n=12)0.0 %
opd_student_otshift-short-ckpt10_m500sub100 / math500greedy3174488.02862999.0 %0.0 %1.0 %66.67 (n=12)88.64 (n=88)98.9 %
opd_student_otshift-short-ckpt10_m500sub100 / math500sample43174460.52202098.0 %1.2 %0.8 %75.32 (n=158)80.17 (n=242)99.2 %
opd_student_otshift-short-ckpt4 / math500greedy317441.296375.8 %22.2 %2.0 %75.71 (n=494)0.00 (n=6)0.0 %
opd_student_otshift-short-ckpt4 / math500sample4317440.265978.8 %20.1 %1.1 %74.04 (n=1995)0.00 (n=5)0.0 %
opd_student_otshift-short-ckpt8 / math500greedy317441.2101298.8 %0.2 %1.0 %75.91 (n=494)66.67 (n=6)66.7 %
opd_student_otshift-short-ckpt8 / math500sample4317440.887098.9 %0.7 %0.4 %75.25 (n=1984)62.50 (n=16)75.0 %
opd_student_otshift_lowlr-ckpt20 / math500greedy317441.087071.8 %25.4 %2.8 %74.75 (n=495)0.00 (n=5)0.0 %
opd_student_otshift_lowlr-ckpt20 / math500sample4317440.470774.0 %23.5 %2.5 %73.83 (n=1991)0.00 (n=9)0.0 %
opd_student_otshift_lowlr-ckpt40 / math500greedy317442.8149497.8 %0.6 %1.6 %77.98 (n=486)50.00 (n=14)50.0 %
opd_student_otshift_lowlr-ckpt40 / math500sample4317440.989098.7 %0.9 %0.4 %74.74 (n=1983)64.71 (n=17)70.6 %
opd_student_otshift_lowlr_m500sub100 / math500greedy31744100.0317440.0 %62.0 %38.0 %0.00 (n=0)54.00 (n=100)62.0 %
opd_student_otshift_lowlr_m500sub100 / math500sample431744100.0317440.8 %61.3 %38.0 %0.00 (n=0)49.00 (n=400)62.0 %
opd_student_r1shift-ckpt20 / math500greedy317440.875898.4 %0.0 %1.6 %73.19 (n=496)0.00 (n=4)0.0 %
opd_student_r1shift-ckpt20 / math500sample4317440.151398.3 %0.2 %1.5 %72.14 (n=1999)0.00 (n=1)0.0 %
opd_student_r1shift-ckpt40 / math500greedy317441.083198.6 %0.0 %1.4 %75.35 (n=495)0.00 (n=5)0.0 %
opd_student_r1shift-ckpt40 / math500sample4317440.154098.8 %0.1 %1.1 %72.62 (n=1998)0.00 (n=2)0.0 %
opd_student_r1shift-ckpt40-full / math500greedy317441.081998.8 %0.0 %1.2 %75.35 (n=495)0.00 (n=5)0.0 %
opd_student_r1shift-ckpt40-full / math500sample4317440.152998.8 %0.1 %1.1 %73.49 (n=1999)0.00 (n=1)0.0 %
opd_student_r1shift_m500sub100 / math500greedy31744100.03174490.0 %0.0 %10.0 %0.00 (n=0)63.00 (n=100)90.0 %
opd_student_r1shift_m500sub100 / math500sample43174496.53067898.2 %0.0 %1.8 %64.29 (n=14)69.95 (n=386)98.2 %
opd_student_raftshift / math500greedy317440.675781.4 %17.4 %1.2 %75.05 (n=497)0.00 (n=3)0.0 %
opd_student_raftshift / math500sample4317440.465381.2 %17.5 %1.2 %72.96 (n=1993)0.00 (n=7)0.0 %
opd_student_raftshift-ckpt20 / math500greedy317440.671586.2 %12.8 %1.0 %76.46 (n=497)0.00 (n=3)0.0 %
opd_student_raftshift-ckpt20 / math500sample4317440.154684.2 %15.0 %0.9 %73.69 (n=1999)0.00 (n=1)0.0 %
opd_student_raftshift-ckpt40 / math500greedy317440.879080.8 %18.4 %0.8 %76.61 (n=496)0.00 (n=4)0.0 %
opd_student_raftshift-ckpt40 / math500sample4317440.261282.2 %16.8 %0.9 %74.79 (n=1995)0.00 (n=5)0.0 %
opd_student_raftshift-ckpt60 / math500greedy317442.0114781.6 %15.8 %2.6 %75.10 (n=490)0.00 (n=10)0.0 %
opd_student_raftshift-ckpt60 / math500sample4317440.159183.2 %15.8 %1.0 %74.11 (n=1997)0.00 (n=3)0.0 %
opd_student_raftshift-ckpt80 / math500greedy317441.083882.2 %16.6 %1.2 %74.55 (n=495)0.00 (n=5)0.0 %
opd_student_raftshift-ckpt80 / math500sample4317440.261582.9 %16.2 %0.9 %73.58 (n=1995)0.00 (n=5)0.0 %
post_teacher / math500greedy3174421.2813473.4 %5.2 %21.4 %83.76 (n=394)0.94 (n=106)0.9 %
post_teacher / math500sample4317440.8382279.8 %17.3 %2.9 %75.81 (n=1984)0.00 (n=16)12.5 %
post_teacher_ot / math500greedy317449.070949.6 %79.4 %11.0 %90.77 (n=455)2.22 (n=45)2.2 %
post_teacher_ot / math500sample4317441.7576212.6 %84.2 %3.3 %90.19 (n=1967)6.06 (n=33)27.3 %
pre_teacher / math500greedy320028.613375.6 %62.6 %31.8 %57.14 (n=357)1.40 (n=143)7.7 %
pre_teacher / math500sample4320018.2105536.7 %41.5 %21.8 %33.82 (n=1635)5.75 (n=365)40.8 %
pre_teacher_ot / math500greedy317443.6160267.8 %19.8 %12.4 %47.10 (n=482)0.00 (n=18)0.0 %
pre_teacher_ot / math500sample4317440.971671.5 %19.1 %9.3 %39.10 (n=1982)0.00 (n=18)0.0 %
raft_teacher / math500greedy317442.8138979.6 %17.2 %3.2 %55.35 (n=486)0.00 (n=14)0.0 %
raft_teacher / math500sample4317441.9110177.5 %19.7 %2.8 %50.94 (n=1961)0.00 (n=39)0.0 %
student_init / math500greedy317440.881184.0 %14.4 %1.6 %76.01 (n=496)0.00 (n=4)0.0 %
student_init / math500sample4317440.156483.7 %14.6 %1.8 %74.42 (n=1998)0.00 (n=2)0.0 %

8.2 Answer-format split, side by side

The same numbers as §8.1, arranged so the direction of the style change is visible. Primary pass only.

benchmarkmodel`Answer:` line`\boxed{}`nonemean out tok
aime24pre_teacher14.0 %61.1 %24.9 %1599
aime24post_teacher21.8 %64.2 %14.1 %13591
aime24student_init58.3 %37.5 %4.2 %1517
aime24opd_student_r1shift93.3 %0.0 %6.7 %29343
aime24opd_student_r1shift-ckpt2096.2 %0.8 %2.9 %1198
aime24opd_student_r1shift-ckpt4098.8 %0.0 %1.2 %1242
aime24opd_student_r1shift-ckpt40-full96.8 %0.0 %3.2 %1467
aime24opd_student_r1shift-ckpt6085.0 %0.0 %15.0 %30993
aime24opd_student_r1shift-ckpt8084.2 %0.0 %15.8 %31744
aime25pre_teacher13.3 %65.6 %21.0 %1523
aime25post_teacher20.6 %72.9 %6.5 %12852
aime25student_init61.8 %36.5 %1.8 %1307
aime25opd_student_r1shift95.8 %0.4 %3.8 %29824
aime25opd_student_r1shift-ckpt40-full99.1 %0.0 %0.9 %988
math500pre_teacher36.7 %41.5 %21.8 %1055
math500post_teacher79.8 %17.3 %2.9 %3822
math500student_init83.7 %14.6 %1.8 %564
math500opd_student_r1shift-ckpt2098.3 %0.2 %1.5 %513
math500opd_student_r1shift-ckpt4098.8 %0.1 %1.1 %540
math500opd_student_r1shift-ckpt40-full98.8 %0.1 %1.1 %529
math500opd_student_r1shift_m500sub10098.2 %0.0 %1.8 %30678

Read the AIME rows first, where the effect is largest. Both teachers lean `\boxed{}` there (61–73 % of samples) — πpre because its own chat template injects *"put your final answer within \boxed{}"*, πpost because the R1 distillation writes that way — and the initial student splits roughly 58/38 between the two formats. The OPD student goes almost purely to the prompt-instructed `Answer:` line and extinguishes `\boxed{}` altogether (37.5 % → 0.8 % by step 20, 0.0 % by step 40). On MATH-500 π_post itself prefers Answer: (79.8 %) and the initial student already does too (83.7 %), so there is less room to move — and the student still moves past both, to 98 %.

The direction is the point. On the benchmark where the two teachers agree on \boxed{}, the student runs the other way. Whatever the token-level log-ratio rewards, it is not "emit π_post's surface form" — the pilot saw the same non-imitation in both of its arms, and it reproduces here with a completely different teacher pair and a 4.7× larger student.

8.3 Extraction on truncated samples — new for this experiment

The pilot could treat truncation as "no answer". This experiment cannot. Two independent reasons, and they point in opposite directions:

  • —π_pre is cut off at 3,200 tokens after it has often already emitted a \boxed{} or an Answer: line, so a large share of its truncated samples still extract — the answer is there, the work around it is not.
  • —The collapsed OPD checkpoints do the opposite in kind and the same in effect — and this is now measured, not predicted. The degenerate policy reaches an answer and then loops on it, so a truncated sample still ends in a gradeable Answer: line. At ckpt-100: 93.2 % of the 221 truncated AIME 2024 samples still extract an answer (and 9.50 pp of them are correct), and on the MATH-500 subset 98.2 % of the 386 truncated sample4 samples extract, scoring 69.95 pp. The collapse is not an answer-extraction failure. The model still solves the problem; it just never stops talking about it. That is the single most useful diagnostic in this report: it separates "the policy was destroyed" from "the policy lost its stop token", and the evidence is for the latter.

Either way, truncation and non-extraction have come apart and both must be reported. Measured now:

unitpasstruncated nextraction successaccuracy on truncated
opd_student_otshift-short-ckpt10_m500sub100 / math500sample424299.2 %80.17 pp
opd_student_otshift-short-ckpt10_m500sub100 / math500greedy8898.9 %88.64 pp
opd_student_r1shift_m500sub100 / math500sample438698.2 %69.95 pp
opd_student_otshift-short-ckpt10 / aime25sample329997.0 %13.13 pp
opd_student_r1shift / aime25sample3222596.0 %8.44 pp
opd_student_otshift_lowlr-ckpt60 / aime24sample3221395.8 %9.39 pp
opd_student_otshift-short-ckpt10 / aime24sample3210693.4 %10.38 pp
opd_student_r1shift / aime24sample3222193.2 %9.50 pp
opd_student_r1shift_m500sub100 / math500greedy10090.0 %63.00 pp
opd_student_r1shift-ckpt60 / aime24sample3223485.0 %6.84 pp
opd_student_r1shift-ckpt80 / aime24sample3224084.2 %6.25 pp
opd_student_otshift_lowlr-ckpt80 / aime24sample3223782.3 %11.39 pp
opd_student_otshift-short-ckpt8 / math500sample41675.0 %62.50 pp
opd_student_otshift_lowlr-ckpt40 / math500sample41770.6 %64.71 pp
opd_student_otshift-short-ckpt8 / math500greedy666.7 %66.67 pp
opd_student_otshift_lowlr / aime24sample3223966.1 %11.30 pp
opd_student_otshift_lowlr / aime25sample3224063.7 %7.50 pp
opd_student_otshift_lowlr_m500sub100 / math500greedy10062.0 %54.00 pp
opd_student_otshift_lowlr_m500sub100 / math500sample440062.0 %49.00 pp
opd_student_otshift_lowlr-ckpt40 / math500greedy1450.0 %50.00 pp
pre_teacher / aime25sample3222546.7 %0.00 pp
opd_student_otshift-short-ckpt8 / aime25sample321145.5 %9.09 pp
pre_teacher / aime24sample3225542.0 %0.00 pp
pre_teacher / math500sample436540.8 %5.75 pp
opd_student_otshift_lowlr-ckpt40 / aime24sample323030.0 %6.67 pp
opd_student_otshift-short-ckpt8 / aime24sample322729.6 %0.00 pp
post_teacher_ot / math500sample43327.3 %6.06 pp
opd_student_otshift_lowlr-ckpt40 / aime25sample32616.7 %0.00 pp
post_teacher / aime25sample323212.5 %0.00 pp
post_teacher / math500sample41612.5 %0.00 pp
post_teacher / aime24sample325111.8 %1.96 pp
post_teacher_ot / aime25sample322129.4 %0.94 pp
post_teacher_ot / aime24sample321827.7 %1.10 pp
pre_teacher / math500greedy1437.7 %1.40 pp
post_teacher_ot / math500greedy452.2 %2.22 pp
raft_teacher / aime24sample321111.8 %0.00 pp
post_teacher / math500greedy1060.9 %0.94 pp
opd_student_otshift-short-ckpt4 / aime24sample32180.0 %0.00 pp
opd_student_otshift_lowlr-ckpt20 / aime24sample32170.0 %0.00 pp
opd_student_r1shift-ckpt20 / aime24sample3220.0 %0.00 pp
opd_student_r1shift-ckpt40 / aime24sample3210.0 %0.00 pp
opd_student_r1shift-ckpt40-full / aime24sample32150.0 %0.00 pp
opd_student_raftshift / aime24sample32230.0 %0.00 pp
opd_student_raftshift-ckpt20 / aime24sample3260.0 %0.00 pp
opd_student_raftshift-ckpt40 / aime24sample3230.0 %0.00 pp
opd_student_raftshift-ckpt60 / aime24sample3260.0 %0.00 pp
opd_student_raftshift-ckpt80 / aime24sample3250.0 %0.00 pp
pre_teacher_ot / aime24sample32620.0 %0.00 pp
student_init / aime24sample32150.0 %0.00 pp
opd_student_otshift-short-ckpt4 / aime25sample32100.0 %0.00 pp
opd_student_otshift_lowlr-ckpt20 / aime25sample3260.0 %0.00 pp
opd_student_r1shift-ckpt40-full / aime25sample3230.0 %0.00 pp
opd_student_raftshift / aime25sample32100.0 %0.00 pp
pre_teacher_ot / aime25sample32460.0 %0.00 pp
raft_teacher / aime25sample32830.0 %0.00 pp
student_init / aime25sample32120.0 %0.00 pp
opd_student_otshift-short-ckpt4 / math500greedy60.0 %0.00 pp
opd_student_otshift-short-ckpt4 / math500sample450.0 %0.00 pp
opd_student_otshift_lowlr-ckpt20 / math500greedy50.0 %0.00 pp
opd_student_otshift_lowlr-ckpt20 / math500sample490.0 %0.00 pp
opd_student_r1shift-ckpt20 / math500greedy40.0 %0.00 pp
opd_student_r1shift-ckpt20 / math500sample410.0 %0.00 pp
opd_student_r1shift-ckpt40 / math500greedy50.0 %0.00 pp
opd_student_r1shift-ckpt40 / math500sample420.0 %0.00 pp
opd_student_r1shift-ckpt40-full / math500greedy50.0 %0.00 pp
opd_student_r1shift-ckpt40-full / math500sample410.0 %0.00 pp
opd_student_raftshift / math500greedy30.0 %0.00 pp
opd_student_raftshift / math500sample470.0 %0.00 pp
opd_student_raftshift-ckpt20 / math500greedy30.0 %0.00 pp
opd_student_raftshift-ckpt20 / math500sample410.0 %0.00 pp
opd_student_raftshift-ckpt40 / math500greedy40.0 %0.00 pp
opd_student_raftshift-ckpt40 / math500sample450.0 %0.00 pp
opd_student_raftshift-ckpt60 / math500greedy100.0 %0.00 pp
opd_student_raftshift-ckpt60 / math500sample430.0 %0.00 pp
opd_student_raftshift-ckpt80 / math500greedy50.0 %0.00 pp
opd_student_raftshift-ckpt80 / math500sample450.0 %0.00 pp
pre_teacher_ot / math500greedy180.0 %0.00 pp
pre_teacher_ot / math500sample4180.0 %0.00 pp
raft_teacher / math500greedy140.0 %0.00 pp
raft_teacher / math500sample4390.0 %0.00 pp
student_init / math500greedy40.0 %0.00 pp
student_init / math500sample420.0 %0.00 pp

Two readings of the table. π_pre: 40–47 % of its truncated samples still extract, against 0–12 % for the two teachers' and the initial student's — so reading "truncated ⇒ no answer" would understate π_pre and, through it, overstate the teacher gain. The collapsed checkpoints: 84–98 % extraction on truncated samples, with real accuracy behind it. Truncation rate alone is therefore a terrible summary of what went wrong here, and this column is what replaces it.

8.4 Sensitivity view — πpost at πpre's cap

πpre is capped at 3,200 tokens by its own trained context, while πpost runs to 31,744. Part of the §5 teacher gain is therefore a budget difference rather than a capability difference. This view costs no new compute: it restricts π_post's stored generations to samples of ≤ 3,200 output tokens and re-scores.

Two readings, because neither alone is honest:

  • —kept-only — accuracy on the π_post samples that fit under 3,200 tokens. This is conditional accuracy and is biased upwards: short samples are the easy problems.
  • —over-cap-as-wrong — every π_post sample longer than 3,200 tokens counted incorrect. This simulates the cap and is a lower bound: a truncated sample sometimes still extracts a correct answer (§8.3).

The true "π_post at 3,200 tokens" lies between them. Both are labelled sensitivity views and neither replaces the §5 primary numbers.

benchmarkpassπ_pre native (pp)π_post native (pp)π_post kept-only (pp)π_post over-cap-as-wrong (pp)samples keptproblems with any kept sample
aime24sample324.69 [1.46, 9.06] (trunc 26.6 %)29.0676.21 [52.01, 99.17]8.33 [2.60, 15.31]9.1 % (87/960)11/30
aime25sample322.40 [0.73, 4.48] (trunc 23.4 %)23.5466.67 [40.00, 91.11]8.75 [2.50, 16.56]9.7 % (93/960)9/30
math500greedy41.20 [36.80, 45.40] (trunc 28.6 %)66.2086.74 [83.00, 90.20]60.20 [56.00, 64.40]69.4 % (347/500)347/500
math500sample428.70 [26.45, 30.95] (trunc 18.2 %)75.2083.60 [80.54, 86.48]60.25 [56.65, 63.75]71.4 % (1427/2000)411/500

The last column is why the kept-only view must not be quoted on AIME: on the 30-problem sets only ~9–11 problems have any π_post sample under 3,200 tokens, and they are the short — i.e. easy — ones. The kept-only row is included for completeness and is not a usable estimate there. On MATH-500, where ~70 % of samples survive the cut and 347–411 of 500 problems are represented, it is informative.

Restated as gains against π_pre's native numbers (paired on the problems each view covers):

benchmarkpassteacher gain, kept-only (pp)nteacher gain, over-cap-as-wrong (pp)n
aime24sample32+64.562 [+41.259, +85.511] p<0.000111+3.646 [-1.042, +9.167] p=0.142630
aime25sample32+61.458 [+35.622, +83.819] p<0.00019+6.354 [+0.625, +13.333] p=0.021630
math500greedy+37.752 [+31.412, +43.804] p<0.0001347+19.000 [+13.600, +24.600] p<0.0001500
math500sample4+50.933 [+47.202, +54.582] p<0.0001411+31.550 [+27.950, +35.050] p<0.0001500

The row that matters is `aime24` / over-cap-as-wrong. It is the pre-registered teacher-gain benchmark, measured with both models on the same token budget and the unmeasurable cases resolved against π_post, and it is the one place where the premise becomes uncertain. MATH-500 keeps a large, budget-independent gain, which is why the co-primary design was worth having.


9. Interpretation vs the pilot

The pilot's finding was that Direct-OPD is a faithful courier of whatever the shift encodes: an SFT shift that encoded narrowing produced a student that narrowed. The obvious follow-up — give it a shift that encodes real capability — is this experiment. The answer is negative, and the interesting part is why. With a teacher pair that demonstrably carried held-out capability, 100 Direct-OPD steps produced a student that is no better on any benchmark at any checkpoint, significantly worse on MATH-500, and — after step ~61 — unable to terminate at all. The pilot's verdict (style transfers, capability does not) survives the strongest test we could give it, and a new failure mode appears alongside it.

What we can say, in decreasing order of confidence.

  1. 1.The premise is no longer the limiting factor, and the result did not follow. The teacher gain is large and unambiguous (+24.375 [+15.417, +34.167] p<0.0001 on AIME 2024). The student's gain, at every checkpoint measured, is not. Whatever blocked transfer here, it is not "the shift had nothing to give" — which was the pilot's explanation and is now excluded.
  2. 2.Surface form moved; capability did not — and the surface form did not move towards π_post. By step 20 the student had abandoned \boxed{} almost entirely in favour of the prompt-instructed Answer: line, while on AIME both teachers lean \boxed{} (§8.2). Accuracy over the same interval is flat on AIME and slightly negative on MATH-500 (§7.4). This is the pilot's finding 5 reproduced with a different teacher pair, a different student family and a 4.7× larger student: a token-level log-ratio objective buys a change of surface form first, and the form it buys is not the post-teacher's. The natural reading is that the log-ratio is largest on tokens where the two teachers disagree about style, and the cheapest way to increase it is to leave the region where they disagree — not to imitate either one.
  3. 3.The failure mode is termination collapse — a lost stop token, not a destroyed policy. It is not a generic "OPD diverged". The chain is legible in the metrics (§6): a reward that is negative on the student's native style → an adaptive KL controller that loosens under sustained negative reward → escape into the degenerate region where πpost ≈ πpre → an Answer: repetition loop that never emits a stop token. It is worse at eval time than in training — at the 31,744-token cap ckpt-100 truncates on 92–100 % of samples despite its training-time partial recovery at the 3,328-token cap — and §8.3 pins down what "truncated" means here: 84–98 % of those truncated samples still extract an answer, and on MATH-500 ~70 pp of them are correct. The model still reasons; it cannot stop. This distinction matters for the fix list below: three of the five candidates target termination specifically.
  4. 4.A capability-bearing shift is not sufficient; distributional proximity may be necessary. The two teachers here are far apart in style (a \boxed{}-instructed math base vs an R1 <think> distillation) and both are far from the student (non-thinking Qwen2.5-Instruct, 4.7× larger). The log-ratio is then dominated by style disagreement, and the capability difference — the thing we wanted — is a small residual underneath it. This is a hypothesis this experiment suggests; it does not test it.

What would need to change, concretely, for a rerun to be informative.

changewhycost
KL floor — clamp the adaptive coefficient below (e.g. ≥ 1.5) instead of letting it ratchet to 1.06the controller loosened the leash precisely in the negative-reward regime that needed it tightenedfree (a config pin)
Length / repetition penalty, or a repetition-aware rewardthe escape region is degenerate repetition; nothing in the objective priced itfree
Response cap raised, or the training cap matched to the eval capthe 3,328-token cap made the collapse cheap to enter and hid its eval-time severitymore GPU-hours per step
A teacher pair closer to the student's own distribution (e.g. both ends Qwen2.5-Instruct-family, or an SFT pair built on the student)makes the log-ratio encode capability rather than stylea teacher-side SFT run
A thinking-mode student (or the student's thinking template at both train and eval time)π_post is a <think>-native model; a non-thinking student can only imitate its surfacefree at eval, config at train

A rerun that changes only one of these is worth more than one that changes all five: the first three test the collapse, the last two test the premise about distance.


10. Limitations

  1. 1.The primary endpoint is measured under a reduced protocol (§2.1(3), §7.1). It is the same estimand with wider intervals, not a different one — but it is not the pre-registered sample count, and a null result at 8 samples/problem is weaker evidence than a null at 32.
  2. 2.MATH-500's ckpt-100 unit is a 100-problem subset, so its interval is roughly √5 wider than the 500-problem baseline's, and it is not the same problem set as the ckpt-20/40 curve points. The curve table says so on every row.
  3. 3.π_pre is context-limited at 3,200 tokens and truncates heavily, so the teacher gain is partly a token-budget difference. §8.4 bounds it but cannot eliminate it: on the pre-registered gate benchmark the bound runs from +3.646 [-1.042, +9.167] p=0.1426 (πpost held to πpre's budget, over-cap counted wrong) up to +24.375 [+15.417, +34.167] p<0.0001 (native protocol). The honest statement of the premise on AIME 2024 is that range, not the headline point. MATH-500 is where the premise survives the correction intact.
  4. 4.n = 30 on AIME. Even at 32 samples/problem the per-problem bootstrap over 30 problems gives intervals ~10 pp wide. MATH-500 is the co-primary precisely because 500 problems give ~4× tighter intervals — which is why losing the full MATH-500 endpoint to cost hurts.
  5. 5.One run, one seed (42), one condition. No RL comparator, no repeat at a different seed, no ablation of the collapse. We cannot separate "Direct-OPD cannot carry this shift" from "this configuration collapsed before it could".
  6. 6.The collapse confounds everything after step ~58. ckpt-60/80/100 are measurements of a degenerate policy. Their accuracies are real, but they answer "what does a collapsed policy score", not "how much capability transferred".
  7. 7.No shift-magnitude diagnostic. The pre-registration folded the reward-sanity check into the P3 smoke to save cost, so there is no weighted-RMS shift magnitude for this pair and therefore no gain-per-unit-shift number comparable to the pilot's.
  8. 8.π_post's AIME units are imported from the pilot, not re-run. Same model revision, same harness, same protocol, verified by provenance.json per unit — but not an independent measurement.
  9. 9.Cost basis. The quoted total is wall-clock occupancy × flavor rate including aborted work; the platform's running_secs basis is a floor (see §11).

11. Cost & reproduction

$332.36 of a $200 cap ($-132.36 unspent), recomputed by summing cost_usd over all 83 rows of ledger/cost_ledger.csv.

phasecostshare
P1$0.000.0 %
P2$7.322.2 %
P2b$16.465.0 %
P3$1.070.3 %
P3b$3.611.1 %
P4$44.4313.4 %
P4b$100.9030.4 %
P5$26.147.9 %
P5b$71.0521.4 %
P6$13.914.2 %
P7$47.4714.3 %
TOTAL$332.36100 %

By hardware: a100x4 $176.73 · a100-large $155.63 · cpu-basic $0.00.

⚠️ Two cost bases. wall-clock occupancy x flavor rate from ledger/costledger.csv, INCLUDING any aborted work — the same (larger) basis the pilot quoted. The platform's own `runningsecs x rate` is a floor, because CANCELED jobs report running_secs = 0 regardless of real occupancy. This report quotes the larger basis, as the pilot did; nothing in the decision record turns on the difference.

⚠️ This ledger is cumulative over run 1, the exp2b extension AND exp3. The cap was raised $200 → $500 by the user on 2026-08-26. Run 1's own phases (P1–P6) total $92.87, inside the original $200 cap; the exp2b extension's *b phases (P2b, P3b, P4b, P5b) add $192.02; exp3's phase (P7) — the shift anatomy, the RAFT pair and its whole eval wave — add $47.47, for $332.36 of $500. §14.4 gives the breakdown, including the cancelled runs booked on the wall-clock basis.

Ledger reconciliation: every eval unit on the Hub has a matching cost-ledger row (units imported from the pilot repo excepted, which cost nothing).

Imported at $0 (no GPU spend expected, provenance.json in each unit): post_teacher/aime24, post_teacher/aime25.

11.1 Reproducing the whole campaign

Everything below is pinned by revision or by sha256, and every table in this subsection is generated from its source at build time (eval_model.MODEL_REGISTRY, ledger/resolved_revisions.json, ledger/staged_driver.json, the files on disk) rather than typed in. A third party with the results repo and this list can rebuild the campaign end to end.

(a) Models — the fixed inputs. Each alias is exactly the --model the harness was launched with.

aliasrolereporevision
pre_teacherteacher_preQwen/Qwen2.5-Math-1.5B4a83ca6e4526a4f2da3aa259ec36c259f66b2ab2
post_teacherteacher_postdeepseek-ai/DeepSeek-R1-Distill-Qwen-1.5Bad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562
student_initstudent_initQwen/Qwen2.5-7B-Instructa09a35458c702b33eeacc393d103063234e8bc28
pre_teacher_otteacher_preQwen/Qwen2.5-1.5B-Instruct989aa7980e4cf806f80c7fef2b1adb7bc71aa306
post_teacher_otteacher_postopen-thoughts/OpenThinker3-1.5B0ee90a38b29bfac8b8b005da9ae32c59e2943785
student_init_qwen3_4bstudent_initQwen/Qwen3-4B1cfa9a7208912126459214e8b04321603b3df60c

(b) Models we trained. Every Direct-OPD run in the study, with the repo and revision its checkpoints live at. A cancelled run published nothing — that is why its evidence in §13.5 is training-only.

runconditionreporevisioncheckpointsstatus
run 1r1shiftcmpatino/Qwen2.5-7B-Instruct-DirectOPD-R1DistillShift-1005819e26cab1f099a9b9b9caa38c6340db50b8c6720, 40, 60, 80, 100trained; checkpoints published (repo root = step 100)
B-shortotshift_shortcmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-10steps8b20748806f63d5b9f05b3239ffdbddc20a62f2a2, 4, 6, 8, 10trained; checkpoints published
B-lowlrotshift_lowlrcmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-lr2e-7-1006ad3add5b065a25c9c003d31ea59b1e191f7e3a720, 40, 60, 80, 100trained; checkpoints published
Botshiftcmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-100——CANCELED at step ~24; no checkpoint was ever published
Cotshift_klfloorcmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-KLfloor-100——CANCELED at step ~24; no checkpoint was ever published
Dotshift_qwen3_4bcmpatino/Qwen3-4B-DirectOPD-OpenThinker3Shift-100——CANCELED at step ~24; no checkpoint was ever published

(c) Datasets and the method itself.

rolereporevision
datacmpatino/direct-opd-sft-deepmath-pilot-data22625ae5db434947195bf862c429cd94504a4809
aime24HuggingFaceH4/aime_20242fe88a2f1091d5048c0f36abc874fb997b3dd99a
aime25yentinglin/aime_20256f71d77b0b89b9dabe07ab466c51df33f514df7f
math500HuggingFaceH4/MATH-5006e4ed1a2a79af7d8630a6b768ec859cb5af4d3be
Direct-OPD (method)https://github.com/BytedTsinghua-SIA/Direct-OPD3a9d6bd37b00a38e7a9b2959239e4631e5324aea (+ code/opd/phase4_seed.patch (carried from the pilot))

(d) Code, by content hash. The driver was staged into the results repo before every launch, so a run's exact driver is recoverable from the revision it was staged at.

filerolesha256 (this working copy)
code/opd/run_opd.shDirect-OPD training driver11f7ff5b78892ccc01cca8bc4fe77e9cbc3d9589be63501f1afb53029c3f1b03
code/eval/eval_model.pyevaluation harness (generation + grading)1720b72e2a8c15e5b0b52a76bd76219d05e582853ae4522610316d566e8d8590
code/eval/aggregate_evals.pycanonical cross-model aggregator355920975a59f7a52fe01fcd402d486e800a537e1f20fbc4ead4184e0a010201
code/report/build_report.pythis report builder2f0b3b701d92e2a384d3991b14a9791cff3754a0c0ad8d6ae9664151d163646f
driver staged at results-repo revision`run_opd.sh` sha256whenwhat changed
a6e5e9924cce0089d0853d557464e1a3a8a51ff3fbeae2b95bb8b0bcb03a6a68bd639f76cdcee6c3f252db72df213204fd2112052026-08-25T10:48:37Zinitial staging
2fabdb13ff2e081ec3d906f938ca1b181849c3af37359fd16a50c01b117f0d6920b9550cb7023841abbec073a1c30edf6a97d02f2026-08-28T12:52:42Zexp2b conditions
b4dfdf37d521db816742bd0fd93af48ec5de18e111f7ff5b78892ccc01cca8bc4fe77e9cbc3d9589be63501f1afb53029c3f1b032026-08-28T14:58:13Zcheckpoint-list driver bug fix: CKPTSTEPS array derived from SAVEFREQ/TOTALTRAININGSTEPS replaces hard-coded '20 40 60 80 100' in the post-train verificatio

(e) The launches themselves. The exact hf jobs run command line of every training launch is stored verbatim, together with the babysitter's per-step tables — which are the only surviving evidence for the three cancelled conditions, since they published no checkpoint. Every one of these files is listed in MANIFEST.json with its sha256 and is uploaded to the results repo under logs/ alongside this bundle:

logs/full_launch_command.txt
logs/launch_smoke_stdout.txt
logs/redirect_launch_commands.txt
logs/exp2b_full_lowlr_steps.tsv
logs/exp2b_full_openthinker_klfloor_steps.tsv
logs/exp2b_full_openthinker_qwen3_4b_steps.tsv
logs/exp2b_full_openthinker_steps.tsv
logs/exp2b_full_short_steps.tsv
logs/exp3_full_raft_steps.tsv
logs/full_steps.tsv

Evaluation launches follow one form (the runbook is code/eval/LAUNCH.md); the reduced-protocol variants differ only in --num-samples/--seeds (AIME) or --limit-problems (MATH-500) and are pushed under their own _sub8 / _m500sub100 aliases so they can never masquerade as canonical units:

sh
# one eval unit
hf jobs uv run -d --image huggingface/trl --secrets HF_TOKEN --flavor a100-large \
  -e HF_HOME=/tmp/hf -e HF_HUB_ENABLE_HF_TRANSFER=1 \
  --label experiment=direct-opd-sft-transfer --label run_id=<run_id> \
  eval_model.py --model <repo-or-alias> --revision <sha> [--subfolder checkpoint-N] \
    --model-alias <alias> --benchmark <aime24|aime25|math500> \
    --output-dir /tmp/exp2-out --skip-if-exists --push

Each training run's own verl artefacts (hydra_config.yaml, hydra_overrides.yaml, run_manifest.json, metrics.jsonl, the gzipped train log) are in the model repo of that run under logs/, and the four 2-step smokes are in the results repo under phase4/smoke/<condition>-<timestamp>/.

(f) The ledger. ledger/cost_ledger.csv — append-only, one row per job, keyed by job id. Every figure in §11 and §14.4 is summed from it at build time, so booking a late job and re-running the build is all it takes to correct the cost of this report; nothing about the cost is frozen into the text.

11.2 Rebuilding this bundle

sh
export HF_HOME=/home/node/local/hf-cache          # HF_TOKEN must be in the environment
/home/node/local/envs/eval-test/bin/python \
  /data/workspaces/direct-opd/exp2-sft-transfer/code/report/build_report.py

# reuse the local cache instead of re-downloading from the Hub:
/home/node/local/envs/eval-test/bin/python \
  /data/workspaces/direct-opd/exp2-sft-transfer/code/report/build_report.py --offline

# the canonical cross-model aggregate (this report reuses its numerics)
/home/node/local/envs/eval-test/bin/python code/eval/aggregate_evals.py --from-hub \
  --results-repo cmpatino/direct-opd-sft-transfer-results \
  --curve-ckpt100-from-primary --output-dir artifacts/eval-agg

The build is idempotent and CPU-only: it downloads (or reuses) every eval unit and every training log, recomputes every statistic from the raw per-problem and per-sample data, regenerates all nine figures, and rewrites aggregate.json, README.md, summary.md, MANIFEST.json and CHECKSUMS.sha256. No number in this report is hand-copied between artifacts, and no number is quoted that does not come from a file listed in MANIFEST.json. Re-running it after the 0 pending unit(s) land is the only action needed to finish the report — every placeholder becomes a measured value and nothing else changes.


12. Artifact inventory

12.1 Evaluation units

aliasbenchmarkstatusproblemssamples/problempassesnote
pre_teacheraime24✅ found3032sample32π_pre baseline (native 3,200-token cap)
pre_teacheraime25✅ found3032sample32π_pre baseline (native 3,200-token cap)
pre_teachermath500✅ found5004greedy, sample4π_pre baseline (native 3,200-token cap)
post_teacheraime24✅ found3032sample32π_post; IMPORTED from the pilot results repo
post_teacheraime25✅ found3032sample32π_post; IMPORTED from the pilot results repo
post_teachermath500✅ found5004greedy, sample4π_post
student_initaime24✅ found3032sample32student baseline, 32 samples
student_initaime25✅ found3032sample32student baseline, 32 samples
student_initmath500✅ found5004greedy, sample4student baseline, greedy + sample4
opd_student_r1shift-ckpt20aime24✅ found308sample32checkpoint curve, 8 samples (seeds 0–7)
opd_student_r1shift-ckpt40aime24✅ found308sample32checkpoint curve, 8 samples (seeds 0–7)
opd_student_r1shift-ckpt60aime24✅ found308sample32checkpoint curve, 8 samples (seeds 0–7)
opd_student_r1shift-ckpt80aime24✅ found308sample32checkpoint curve, 8 samples (seeds 0–7)
opd_student_r1shift-ckpt20math500✅ found5004greedy, sample4checkpoint curve, full 500-problem protocol
opd_student_r1shift-ckpt40math500✅ found5004greedy, sample4checkpoint curve, full 500-problem protocol
opd_student_r1shiftaime24✅ found308sample32PRIMARY endpoint, REDUCED protocol (8 samples, seeds 0–7)
opd_student_r1shiftaime25✅ found308sample32PRIMARY endpoint, REDUCED protocol (8 samples, seeds 0–7)
opd_student_r1shift_m500sub100math500✅ found1004greedy, sample4PRIMARY endpoint, REDUCED protocol (100-problem subset)
opd_student_r1shift-ckpt40-fullaime24✅ found3032sample32A1: run-1 ckpt-40 at the FULL protocol (32 samples)
opd_student_r1shift-ckpt40-fullaime25✅ found3032sample32A1: run-1 ckpt-40 at the FULL protocol (32 samples)
opd_student_r1shift-ckpt40-fullmath500✅ found5004greedy, sample4A1: run-1 ckpt-40 at the FULL protocol (500 problems)
pre_teacher_otaime24✅ found3032sample32exp2b pi_pre (Qwen2.5-1.5B-Instruct), full 31,744 cap
pre_teacher_otaime25✅ found3032sample32exp2b pi_pre, full 31,744 cap
pre_teacher_otmath500✅ found5004greedy, sample4exp2b pi_pre, full 31,744 cap
post_teacher_otaime24✅ found3032sample32exp2b pi_post (OpenThinker3-1.5B), full 31,744 cap
post_teacher_otaime25✅ found3032sample32exp2b pi_post, full 31,744 cap
post_teacher_otmath500✅ found5004greedy, sample4exp2b pi_post, full 31,744 cap
opd_student_otshift-short-ckpt4aime24✅ found3032sample32B-short step-4 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt4aime25✅ found3032sample32B-short step-4 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt4math500✅ found5004greedy, sample4B-short step-4 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt8aime24✅ found3032sample32B-short step-8 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt8aime25✅ found3032sample32B-short step-8 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt8math500✅ found5004greedy, sample4B-short step-8 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt10aime24✅ found308sample32B-short step-10 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt10aime25✅ found308sample32B-short step-10 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift-short-ckpt10math500✅ found1004greedy, sample4B-short step-10 endpoint, full protocol (reduced if non-terminating) — measured under the REDUCED protocol as opd_student_otshift-short-ckpt10_m500sub100
opd_student_otshift_lowlr-ckpt20aime24✅ found3032sample32B-lowlr step-20 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlr-ckpt20aime25✅ found3032sample32B-lowlr step-20 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlr-ckpt20math500✅ found5004greedy, sample4B-lowlr step-20 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlr-ckpt40aime24✅ found3032sample32B-lowlr step-40 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlr-ckpt40aime25✅ found3032sample32B-lowlr step-40 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlr-ckpt40math500✅ found5004greedy, sample4B-lowlr step-40 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlr-ckpt60aime24✅ found308sample32B-lowlr checkpoint curve, 8 samples (seeds 0-7)
opd_student_otshift_lowlr-ckpt80aime24✅ found308sample32B-lowlr checkpoint curve, 8 samples (seeds 0-7)
opd_student_otshift_lowlraime24✅ found308sample32B-lowlr step-100 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlraime25✅ found308sample32B-lowlr step-100 endpoint, full protocol (reduced if non-terminating)
opd_student_otshift_lowlrmath500✅ found1004greedy, sample4B-lowlr step-100 endpoint, full protocol (reduced if non-terminating) — measured under the REDUCED protocol as opd_student_otshift_lowlr_m500sub100
raft_teachermath500✅ found5004greedy, sample4exp3 pi_post^RAFT, gate G2 primary (full 500-problem protocol)
raft_teacheraime24✅ found3032sample32exp3 pi_post^RAFT, gate G2 secondary
raft_teacheraime25✅ found3032sample32exp3 pi_post^RAFT, gate G2 secondary
raft_teacherteacher_eval✅ found5124greedy, sample4exp3 pi_post^RAFT, in-distribution probe (pilot pool held-out 512-prompt split)
pre_teacher_otteacher_eval✅ found5124greedy, sample4exp3 pi_pre baseline on the same in-distribution probe
opd_student_raftshift-ckpt20aime24✅ found308sample32RAFT step-20 checkpoint curve, 8 samples (seeds 0-7)
opd_student_raftshift-ckpt20math500✅ found5004greedy, sample4RAFT step-20 endpoint, full protocol
opd_student_raftshift-ckpt40aime24✅ found308sample32RAFT step-40 checkpoint curve, 8 samples (seeds 0-7)
opd_student_raftshift-ckpt40math500✅ found5004greedy, sample4RAFT step-40 endpoint, full protocol
opd_student_raftshift-ckpt60aime24✅ found308sample32RAFT step-60 checkpoint curve, 8 samples (seeds 0-7)
opd_student_raftshift-ckpt60math500✅ found5004greedy, sample4RAFT step-60 endpoint, full protocol
opd_student_raftshift-ckpt80aime24✅ found308sample32RAFT step-80 checkpoint curve, 8 samples (seeds 0-7)
opd_student_raftshift-ckpt80math500✅ found5004greedy, sample4RAFT step-80 endpoint, full protocol
opd_student_raftshiftaime24✅ found3032sample32RAFT step-100 endpoint, full protocol
opd_student_raftshiftaime25✅ found3032sample32RAFT step-100 endpoint, full protocol
opd_student_raftshiftmath500✅ found5004greedy, sample4RAFT step-100 endpoint, full protocol
opd_student_r1shift-ckpt60math500⛔ skipped———cost-forced skip: re-priced $115 high, over the $8/checkpoint gate
opd_student_r1shift-ckpt80math500⛔ skipped———cost-forced skip: re-priced $124 high, over the $8/checkpoint gate
opd_student_r1shiftmath500⛔ skipped———superseded by the 100-problem subset unit: full 500-problem protocol re-priced $100/$118 high
opd_student_otshiftaime24⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshiftaime25⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshiftmath500⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift-ckpt20aime24⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift-ckpt40aime24⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift-ckpt60aime24⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift-ckpt80aime24⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift-ckpt20math500⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift-ckpt40math500⛔ canceled / dropped———condition B CANCELED at step ~24 after termination collapse — collapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would only replicate run 1's collapsed-checkpoint evaluations; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klflooraime24⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klflooraime25⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloormath500⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloor-ckpt20aime24⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloor-ckpt40aime24⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloor-ckpt60aime24⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloor-ckpt80aime24⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloor-ckpt20math500⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_klfloor-ckpt40math500⛔ canceled / dropped———condition C CANCELED at step ~24 after termination collapse — collapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4baime24⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4baime25⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4bmath500⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4b-ckpt20aime24⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4b-ckpt40aime24⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4b-ckpt60aime24⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4b-ckpt80aime24⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4b-ckpt20math500⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_otshift_qwen3_4b-ckpt40math500⛔ canceled / dropped———condition D CANCELED at step ~24 after termination collapse — collapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short because thinking-mode evaluations cost ~$65/unit; no checkpoint was ever published, so no evaluation unit can exist
opd_student_r1shift-fullaime24⛔ canceled / dropped———tier A2 DROPPED by the user on 2026-08-26 (PREREGISTRATION E4) to keep condition D inside the $500 cap; run 1's ckpt-100 keeps its logged reduced-protocol asterisk
student_init_qwen3_4baime24⛔ canceled / dropped———condition-D baseline: D was trimmed to aime24 + math500 on 2026-08-26 (E4) and then CANCELED outright on 2026-08-28 after collapsing, so its baseline was never launched (thinking-mode evaluations were priced at ~$35-40 per benchmark)
opd_student_r1shift-fullaime25⛔ canceled / dropped———tier A2 DROPPED by the user on 2026-08-26 (PREREGISTRATION E4) to keep condition D inside the $500 cap; run 1's ckpt-100 keeps its logged reduced-protocol asterisk
student_init_qwen3_4baime25⛔ canceled / dropped———condition-D baseline: D was trimmed to aime24 + math500 on 2026-08-26 (E4) and then CANCELED outright on 2026-08-28 after collapsing, so its baseline was never launched (thinking-mode evaluations were priced at ~$35-40 per benchmark)
opd_student_r1shift-fullmath500⛔ canceled / dropped———tier A2 DROPPED by the user on 2026-08-26 (PREREGISTRATION E4) to keep condition D inside the $500 cap; run 1's ckpt-100 keeps its logged reduced-protocol asterisk
student_init_qwen3_4bmath500⛔ canceled / dropped———condition-D baseline: D was trimmed to aime24 + math500 on 2026-08-26 (E4) and then CANCELED outright on 2026-08-28 after collapsing, so its baseline was never launched (thinking-mode evaluations were priced at ~$35-40 per benchmark)
opd_student_otshift-short-ckpt10_m500sub100math500✅ found1004greedy, sample4exp2b REDUCED-protocol fallback unit (the checkpoint did not terminate at the 31,744 cap)
opd_student_otshift_lowlr_m500sub100math500✅ found1004greedy, sample4exp2b REDUCED-protocol fallback unit (the checkpoint did not terminate at the 31,744 cap)

Sample-level alignment audit (aggregate_evals.audit_alignment, run per comparison group): OK.

  • —group full: ok=True
  • —group curve8: ok=True
  • —group sub100: ok=True
  • —group exp2b-full: ok=True — aime24: grids differ but are NESTED {'opdstudentotshift-short-ckpt10': 240, 'opdstudentotshift-short-ckpt4': 960, 'opdstudentotshift-short-ckpt8': 960, 'opdstudentotshiftlowlr': 240, 'opdstudentotshiftlowlr-ckpt20': 960, 'opdstudentotshiftlowlr-ckpt40': 960, 'opdstudentr1shift-ckpt40-full': 960, 'postteacherot': 960, 'preteacherot': 960} — expected when diagnostic curve units (8 samples/problem) sit beside primary units; aime25: grids differ but are NESTED {'opdstudentotshift-short-ckpt10': 240, 'opdstudentotshift-short-ckpt4': 960, 'opdstudentotshift-short-ckpt8': 960, 'opdstudentotshiftlowlr': 240, 'opdstudentotshiftlowlr-ckpt20': 960, 'opdstudentotshiftlowlr-ckpt40': 960, 'opdstudentr1shift-ckpt40-full': 960, 'postteacherot': 960, 'preteacherot': 960} — expected when diagnostic curve units (8 samples/problem) sit beside primary units
  • —group unexpected: ok=True
  • —group exp2b-curve8: ok=True
  • —group exp3-full: ok=True
  • —group exp3-curve8: ok=True

12.2 This bundle

aggregate.json        every statistic in this report, machine-readable
summary.md            the same, rendered as tables (also pushed to evals/aggregate/)
README.md             this file
figures/f1..f5*.png   Phase-4 training figures, reused verbatim from artifacts/figures/
figures/f6*.png       student gain vs OPD step, paired CIs, collapse band
figures/f7*.png       teacher gain vs student gain
MANIFEST.json         every input file, its sha256 and its repo revision
CHECKSUMS.sha256      sha256 of every file in this bundle

MANIFEST.json lists 186 input files. Large inputs that already live in the results repo (the generations.parquet of every unit) are referenced by repo + revision + sha256, not duplicated here.


13. Extension exp2b — a second teacher pair, three cancelled conditions, two redirects

Pre-registered PREREGISTRATION.md sections E1-E3 (2026-08-26), extended by §E5 (2026-08-28). Budget cap raised to $500 cumulative. Sections 0–12 above are run 1 and are unaffected by anything here.

Status: 32 of 32 live exp2b units are on the Hub, and a further 33 pre-registered units are canceled or dropped — they can never exist, because the runs that would have produced them were stopped before their first checkpoint save (§13.5) or the tier was dropped by the user. Every cell below is a measured value; nothing here is an estimate and nothing is outstanding.

13.1 The new teacher pair, and why it needs no cap deviation

  • —π_pre = Qwen/Qwen2.5-1.5B-Instruct @ 989aa7980e4cf806f80c7fef2b1adb7bc71aa306
  • —πpost = `open-thoughts/OpenThinker3-1.5B @ 0ee90a38b29bfac8b8b005da9ae32c59e2943785` — **SFT-only** (7 epochs on OpenThoughts3-1.2M) from that same πpre, i.e. the same family and post-training lineage as the 7B student.
  • —shift key openthinker3_sft

Both models have maxpositionembeddings 32,768 and the longest prompt under their (identical) chat template is 872 tokens, so 872 + 31,744 = 32,616 fits and neither carries a cap deviation — unlike run 1's pi_pre.

modelmax_position_embeddingslongest promptprompt + 31,744cap deviation
pre_teacher_ot32,768872 (MATH-500)32,616none
post_teacher_ot32,768872 (MATH-500)32,616none
(run 1) pre_teacher4,096871 (MATH-500)32,615 — does not fit--max-tokens 3200, on every unit

This is the single biggest measurement difference between run 1 and exp2b: run 1's AIME teacher gain shrank from +24.38 pp to +3.65 [−1.04, +9.17] once πpost was held to πpre's 3,200-token budget (§8's cap-sensitivity view). The exp2b pair has no such confound — both models are measured at the same, full, pre-registered cap.

13.2 Teacher gains for the new pair — the exp2b premise

benchmarkpassπ_pre (pp)π_post (pp)gain (pp)95 % CIp
aime24sample322.18853.958+51.771[+38.646, +64.375]p<0.0001
aime25sample320.41743.021+42.604[+28.539, +56.667]p<0.0001
math500greedy45.40082.800+37.400[+32.400, +42.400]p<0.0001
math500sample438.75088.800+50.050[+46.650, +53.400]p<0.0001

TEACHER-GAIN GATE (AIME 2024): PASS — +51.771 pp [+38.646, +64.375].

PREREGISTRATION E3: paired post-pre on AIME24 must be > 0 with a 95 % CI excluding 0 (expected ~ +49 pp from the published 3.0 -> 52.0)

These are the paired gains the gate is stated on. They are what makes the rest of this section a result rather than a null: the shift this pair encodes is real and large, so when four separate students fail to inherit any of it, the explanation cannot be that the shift had nothing to give.

13.3 A1 — run-1 ckpt-40 at the FULL protocol

ckpt-40 is the last pre-collapse checkpoint of run 1 and it terminates (0.4 % truncation). It was only ever measured on the 8-sample curve protocol; at the full protocol it becomes the study's one ckpt-vs-init comparison that needs no caveat at all — same 32 seeds, same 500 problems, on both sides.

Alias opd_student_r1shift-ckpt40-full (a separate directory, so the published 8-sample curve units at opd_student_r1shift-ckpt40/* are untouched), baseline student_init.

benchmarkpasssamples/problemfull protocol?ckpt-40 (pp)student_init (pp)gain (pp)n
aime24sample3232yes10.00012.188-2.188 [-4.583, -0.310] p=0.015230
aime25sample3232yes7.8126.979+0.833 [-1.667, +3.750] p=0.522630
math500greedy4yes74.60075.400-0.800 [-4.200, +2.600] p=0.6040500
math500sample44yes73.45074.350-0.900 [-2.650, +0.800] p=0.2972500

13.4 A2 — run-1 ckpt-100 at the FULL protocol: DROPPED

DROPPED by the user on 2026-08-26 (PREREGISTRATION E4) to keep condition D inside the $500 cap; run 1's ckpt-100 keeps its logged reduced-protocol asterisk. Its three roster rows read canceled, never pending.

13.5 Conditions B, C and D — CANCELED after termination collapse (training-only)

All three were launched on 2026-08-28 and all three collapsed. Their pre-registered first checkpoint save is at step 20, and all three crossed the collapse-watch threshold before it, so none of them published a policy that could be evaluated. They were cancelled at step ~22–25 by user decision (PREREGISTRATION §E5) and their GPU time was booked on the wall-clock basis. What follows is therefore training-only evidence: no accuracy number for B, C or D exists anywhere in this report, and none is estimated.

conditionteacher pairstudentlrKL regimeresp capsteps reachedlength runawaycollapse onsetcheckpointshypothesisverdict
B — otshiftopenthinker3_sftQwen2.5-7B-Instruct1e-6adaptive [0.5, 2.5], eps 0.013,32825 / 1001012noneH2UNTESTED
C — otshift_klflooropenthinker3_sftQwen2.5-7B-Instruct1e-6CONSTANT 2.5 (ADAPTIVEKLLOSSMINCOEF = MAX = 2.5)3,32824 / 1001012noneH3REFUTED
D — otshiftqwen34bopenthinker3_sftQwen3-4B (thinking)1e-6adaptive [0.5, 2.5], eps 0.014,09622 / 1001717noneH4REFUTED

Per-step snapshots at steps 1 / 5 / 10 / 15 / 20, read from the babysitter's per-step tables (the only surviving record of a cancelled run — nothing was ever uploaded):

conditionstepresponse len (mean)clip ratioactor entropyweighted shift rewardKL coefgrad norm
run 113390.0000.156-0.00782.4754.24
run 153730.0000.139-0.00602.3771.83
run 1103780.0000.162-0.00532.2611.59
run 1154210.0000.187-0.00472.1501.40
run 1203990.0000.171-0.00412.0451.53
run 1405190.0000.217-0.00341.6720.92
run 16011810.2190.134-0.00061.3680.85
run 16124890.6800.069-0.00031.3540.28
run 18033110.9920.094-0.00051.1190.46
run 110026140.6990.667+0.00261.1182.10
B13390.0000.156-0.03542.4753.43
B54330.0000.107-0.02842.37726.53
B108820.0980.075-0.01502.26121.30
B1230590.8160.030-0.00372.2162.40
B1533281.0000.032-0.00312.1500.60
B2033281.0000.080-0.00292.0451.24
B2533281.0000.088-0.00371.9450.88
C13390.0000.156-0.03542.5003.43
C54450.0000.103-0.02732.50023.99
C1010820.1640.065-0.01202.50018.38
C1231180.8360.031-0.00362.5001.91
C1533281.0000.033-0.00322.5000.69
C2033281.0000.084-0.00332.5001.45
C2433281.0000.077-0.00432.5000.92
D117070.0510.319-0.01432.4754.42
D520030.0550.241-0.00952.3773.76
D1024690.1050.289-0.00582.2612.18
D1526370.1520.305-0.00432.1502.27
D1737240.6450.311-0.00352.1073.82
D2040820.9770.336+0.00262.1282.66
D2240890.9920.389-0.00112.1281.58
B-short13390.0000.156-0.03542.4753.43
B-short54390.0000.109-0.02812.3779.40
B-short1015210.2890.048-0.00852.26114.82
B-lowlr13390.0000.156-0.03542.4753.43
B-lowlr54030.0000.115-0.03042.3772.87
B-lowlr104190.0040.122-0.02982.2612.27
B-lowlr154860.0000.114-0.02572.1503.75
B-lowlr204620.0000.119-0.02742.04518.77
B-lowlr406130.0080.129-0.02271.6727.41
B-lowlr427870.0350.154-0.01951.63910.89
B-lowlr4824550.5040.051-0.00591.5436.25
B-lowlr6033281.0000.041-0.00411.3680.51
B-lowlr8033180.9960.112-0.00441.1191.15
B-lowlr10033281.0000.169-0.00380.9151.36

Grad-norm peaks and the KL controller. The two spikes that precede each collapse are the clearest early-warning signal in the whole study.

conditiongrad-norm peak (step)spikes ≥ 10KL coef first → lastconstant?weighted shift reward first → laststeps with a negative reward
run 14.2 (step 1)none2.475 → 1.118no-0.0078 → +0.002690 / 100
B26.5 (step 5)27@5, 13@9, 21@102.475 → 1.945no-0.0354 → -0.003725 / 25
C24.0 (step 5)24@5, 11@6, 12@9, 18@102.500 → 2.500yes-0.0354 → -0.004324 / 24
D4.4 (step 1)none2.475 → 2.128no-0.0143 → -0.001119 / 22
B-short15.9 (step 9)10@4, 16@9, 15@102.475 → 2.261no-0.0354 → -0.008510 / 10
B-lowlr49.8 (step 21)13@19, 19@20, 50@21, 18@22, 12@23, 10@412.475 → 0.915no-0.0354 → -0.0038100 / 100

[image]

[image]

The two hypotheses these runs did settle, and they settled them by refutation.

  • —H2 — UNTESTED (condition B): Collapsed at step 12; the first checkpoint save is at step 20, i.e. after the collapse, so no pre-collapse policy exists to evaluate. Recorded at the time: UNTESTED — the run never produced a pre-collapse checkpoint and was cancelled before any evaluation. H2 is carried by the B-short and B-lowlr redirects instead.
  • —H3 — REFUTED (condition C): The KL coefficient was constant at 2.5 for all 24 recorded steps and the clip ratio still crossed 0.5 at step 12. Recorded at the time: REFUTED — the KL coefficient was 2.500 at every step and the run still collapsed, on the same step as B. A constant KL coefficient does not prevent termination collapse.
  • —H4 — REFUTED (condition D): H4's premise was that the shift rewards the thinking student's native distribution (weightedrewardmean > 0 at step 1); the measured value is -0.01427, and the run collapsed at step 17 anyway. Recorded at the time: REFUTED — H4's own premise failed first: the step-1 weighted shift reward on the THINKING student was NEGATIVE, the same signature run 1 showed, and the run then collapsed like every other. The thinking student delayed the sink by ~5 steps; it did not avoid it.

Note what H3's refutation costs the fix list: the KL brake at its maximum does not prevent the sink. C held its coefficient at 2.5 at every single step — the driver's kl_controller clamps max(min, min(max, x)), and min == max pins it — and it collapsed on the same step as B, with a trajectory that matches B's to three decimal places for the first nine steps. A stronger anchor is not the answer, because the anchor was already at its strongest.

13.6 B-short — OpenThinker3 shift, 10 steps, save every 2

  • —repo cmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-10steps @ 8b20748806f63d5b (pinned (training COMPLETE 2026-08-28, job 6a91a23c))
  • —baseline student_init · teacher pair post_teacher_ot − pre_teacher_ot · lr 1e-6 · KL adaptive [0.5, 2.5], eps 0.01
  • —Captures the policies B and C never saved. Trajectory reproduces B/C exactly (same seed): clip 0 through step 8, then response mean 474 -> 1,521 and clip 0.01 -> 0.29 at step 10.
  • —Did ANY held-out capability move before the termination sink? ckpt-4 and ckpt-8 are clean pre-collapse policies (clip_ratio 0 through step 8); ckpt-10 is the onset step itself.

Training record.

10 of 10 pre-registered steps recorded (source: metrics.jsonl (model repo)). Collapse onset (first clip_ratio > 0.5): none; length-runaway marker (mean rollout length past 2x its step-1 value): step 10; peak clip ratio 0.289. Grad-norm peak 15.9 at step 9; KL coefficient 2.475 → 2.261; weighted shift reward -0.0354 → -0.0085, negative on 10 of 10 steps.

Note the two markers disagree, and the disagreement is the point: by the pre-registered collapse-watch rule (clip_ratio > 0.5) this run never collapsed, but the mean rollout length had already more than doubled by step 10 (peak clip 0.289). The onset signature is present; the run simply ended before the clip ratio caught up.
conditionstepresponse len (mean)clip ratioactor entropyweighted shift rewardKL coefgrad norm
B-short13390.0000.156-0.03542.4753.43
B-short54390.0000.109-0.02812.3779.40
B-short1015210.2890.048-0.00852.26114.82

Held-out gains, per evaluated checkpoint (paired per-problem bootstrap vs student_init; the protocol is detected from each unit, never assumed)

stepunitbenchmarkprotocolpairingckpt (pp)baseline (pp)gain (pp)n
4opd_student_otshift-short-ckpt4aime24fullseeds 0–31 (PRIMARY)11.4612.19-0.729 [-2.292, +0.521] p=0.266230
4opd_student_otshift-short-ckpt4aime25fullseeds 0–31 (PRIMARY)7.406.98+0.417 [-1.042, +1.875] p=0.554230
4opd_student_otshift-short-ckpt4math500fullgreedy74.8075.40-0.600 [-3.200, +2.000] p=0.5996500
4opd_student_otshift-short-ckpt4math500fullsample4 (PRIMARY)73.8574.35-0.500 [-2.050, +1.050] p=0.5036500
8opd_student_otshift-short-ckpt8aime24fullseeds 0–31 (PRIMARY)11.0412.19-1.146 [-3.229, +0.625] p=0.200230
8opd_student_otshift-short-ckpt8aime25fullseeds 0–31 (PRIMARY)7.816.98+0.833 [-0.833, +2.500] p=0.299230
8opd_student_otshift-short-ckpt8math500fullgreedy75.8075.40+0.400 [-2.600, +3.400] p=0.7314500
8opd_student_otshift-short-ckpt8math500fullsample4 (PRIMARY)75.1574.35+0.800 [-0.950, +2.550] p=0.3598500
10opd_student_otshift-short-ckpt10aime24reducedseeds 0–7 (PRIMARY)10.4213.75-3.333 [-7.500, +0.417] p=0.072230
10opd_student_otshift-short-ckpt10aime24reducedvs the full 32-sample baseline (unpaired-in-samples)10.4212.19-1.771 [-4.375, +0.521] p=0.117030
10opd_student_otshift-short-ckpt10aime25reducedseeds 0–7 (PRIMARY)10.006.25+3.750 [-0.833, +8.750] p=0.104830
10opd_student_otshift-short-ckpt10aime25reducedvs the full 32-sample baseline (unpaired-in-samples)10.006.98+3.021 [-0.625, +7.187] p=0.102230
10opd_student_otshift-short-ckpt10_m500sub100math500reduced-100-problem-subsetgreedy86.0084.00+2.000 [-5.000, +9.000] p=0.4930100
10opd_student_otshift-short-ckpt10_m500sub100math500reduced-100-problem-subsetsample4 (PRIMARY)78.2577.75+0.500 [-3.750, +4.750] p=0.7828100

Truncation at the 31,744-token evaluation cap, per checkpoint: step 4 aime24 1.9 %; step 4 aime25 1.0 %; step 4 math500 0.2 %; step 8 aime24 2.8 %; step 8 aime25 1.1 %; step 8 math500 0.8 %; step 10 aime24 44.2 %; step 10 aime25 41.2 %; step 10 math500 60.5 %.

⚠ Read the high-truncation rows as termination artifacts, not as capability being destroyed. step 10 on math500 is 60.5 % truncated and scores +0.50 pp. A policy that never emits a stop token produces a 31,744-token sample with no final answer in it, and the grader scores that as wrong however much the model knows — B-short's ckpt-10 is the control for exactly this: it is 41–60 % truncated and its accuracy is nonetheless preserved, because its answers still appear before the loop begins. The deeper into the sink a checkpoint is, the more of its samples never reach an answer at all.

Transfer ratio at the terminal checkpoint (student gain ÷ the OpenThinker3 pair's teacher gain, five guards)

benchmarkteacher gainstudent gainnumerator CI excludes 0ratioreason
aime24+51.771 [+38.646, +64.375]-3.333 [-7.500, +0.417]Falsenullguard: student_gain = -3.333 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BOTH gains are positive; a non-positive nu
aime25+42.604 [+28.539, +56.667]+3.750 [-0.833, +8.750]Falsenullguard (exp2b, new): student_gain = 3.750 pp has CI95 [-0.833, 8.750] pp which INCLUDES zero — the numerator is statistically indistinguishable from 'n
math500+50.050 [+46.650, +53.400]+0.500 [-3.750, +4.750]Falsenullguard (exp2b, new): student_gain = 0.500 pp has CI95 [-3.750, 4.750] pp which INCLUDES zero — the numerator is statistically indistinguishable from 'n

H2 (pre-collapse form) verdict: NOT SUPPORTED

  • —gain_on_at_least_one_benchmark — FAILED: no benchmark shows a positive paired gain whose 95 % CI excludes 0

13.7 B-lowlr — OpenThinker3 shift at OPTIM_LR 2e-7, 100 steps

  • —repo cmpatino/Qwen2.5-7B-Instruct-DirectOPD-OpenThinker3Shift-lr2e-7-100 @ 6ad3add5b065a25c (resolved from `main` at build time (training COMPLETE 2026-08-28, job 6a919e1c, 100/100 steps))
  • —baseline student_init · teacher pair post_teacher_ot − pre_teacher_ot · lr 2e-7 (logged DEVIATION from the pinned 1e-6) · KL adaptive [0.5, 2.5], eps 0.01
  • —The 'can it be made to work' run: smaller steps so the policy can follow the gentle shift gradient without the grad-norm-spike-driven jump into the repetition sink (spikes of 26.5 and 21 preceded B's and C's collapses).
  • —H5: with lr 2e-7 the run does NOT collapse (clip_ratio < 0.5 through step 100) AND ckpt-100 shows a paired gain on at least one held-out benchmark.

Training record.

100 of 100 pre-registered steps recorded (source: metrics.jsonl (model repo)). Collapse onset (first clip_ratio > 0.5): 48; length-runaway marker (mean rollout length past 2x its step-1 value): step 42; peak clip ratio 1.000. Grad-norm peak 49.8 at step 21; KL coefficient 2.475 → 0.915; weighted shift reward -0.0354 → -0.0038, negative on 100 of 100 steps.

conditionstepresponse len (mean)clip ratioactor entropyweighted shift rewardKL coefgrad norm
B-lowlr13390.0000.156-0.03542.4753.43
B-lowlr54030.0000.115-0.03042.3772.87
B-lowlr104190.0040.122-0.02982.2612.27
B-lowlr154860.0000.114-0.02572.1503.75
B-lowlr204620.0000.119-0.02742.04518.77
B-lowlr406130.0080.129-0.02271.6727.41
B-lowlr427870.0350.154-0.01951.63910.89
B-lowlr4824550.5040.051-0.00591.5436.25
B-lowlr6033281.0000.041-0.00411.3680.51
B-lowlr8033180.9960.112-0.00441.1191.15
B-lowlr10033281.0000.169-0.00380.9151.36

Held-out gains, per evaluated checkpoint (paired per-problem bootstrap vs student_init; the protocol is detected from each unit, never assumed)

stepunitbenchmarkprotocolpairingckpt (pp)baseline (pp)gain (pp)n
20opd_student_otshift_lowlr-ckpt20aime24fullseeds 0–31 (PRIMARY)12.0812.19-0.104 [-2.396, +1.875] p=0.905030
20opd_student_otshift_lowlr-ckpt20aime25fullseeds 0–31 (PRIMARY)6.886.98-0.104 [-1.354, +1.042] p=0.805830
20opd_student_otshift_lowlr-ckpt20math500fullgreedy74.0075.40-1.400 [-3.800, +1.000] p=0.2108500
20opd_student_otshift_lowlr-ckpt20math500fullsample4 (PRIMARY)73.5074.35-0.850 [-2.300, +0.550] p=0.2170500
40opd_student_otshift_lowlr-ckpt40aime24fullseeds 0–31 (PRIMARY)11.4612.19-0.729 [-2.917, +0.833] p=0.435030
40opd_student_otshift_lowlr-ckpt40aime25fullseeds 0–31 (PRIMARY)7.406.98+0.417 [-1.979, +2.396] p=0.622030
40opd_student_otshift_lowlr-ckpt40math500fullgreedy77.2075.40+1.800 [-1.200, +4.800] p=0.2088500
40opd_student_otshift_lowlr-ckpt40math500fullsample4 (PRIMARY)74.6574.35+0.300 [-1.350, +2.000] p=0.7118500
60opd_student_otshift_lowlr-ckpt60aime24reducedseeds 0–7 (PRIMARY)9.1713.75-4.583 [-10.000, +0.000] p=0.033430
60opd_student_otshift_lowlr-ckpt60aime24reducedvs the full 32-sample baseline (unpaired-in-samples)9.1712.19-3.021 [-7.083, -0.104] p=0.034230
80opd_student_otshift_lowlr-ckpt80aime24reducedseeds 0–7 (PRIMARY)11.2513.75-2.500 [-6.667, +1.250] p=0.167030
80opd_student_otshift_lowlr-ckpt80aime24reducedvs the full 32-sample baseline (unpaired-in-samples)11.2512.19-0.938 [-4.271, +2.708] p=0.555830
100opd_student_otshift_lowlraime24reducedseeds 0–7 (PRIMARY)11.2513.75-2.500 [-7.917, +2.500] p=0.302430
100opd_student_otshift_lowlraime24reducedvs the full 32-sample baseline (unpaired-in-samples)11.2512.19-0.938 [-5.312, +2.917] p=0.657830
100opd_student_otshift_lowlraime25reducedseeds 0–7 (PRIMARY)7.506.25+1.250 [-2.500, +4.583] p=0.405830
100opd_student_otshift_lowlraime25reducedvs the full 32-sample baseline (unpaired-in-samples)7.506.98+0.521 [-2.604, +3.542] p=0.689230
100opd_student_otshift_lowlr_m500sub100math500reduced-100-problem-subsetgreedy54.0084.00-30.000 [-40.000, -20.000] p<0.0001100
100opd_student_otshift_lowlr_m500sub100math500reduced-100-problem-subsetsample4 (PRIMARY)49.0077.75-28.750 [-35.000, -22.744] p<0.0001100

Truncation at the 31,744-token evaluation cap, per checkpoint: step 20 aime24 1.8 %; step 20 aime25 0.6 %; step 20 math500 0.4 %; step 40 aime24 3.1 %; step 40 aime25 0.6 %; step 40 math500 0.9 %; step 60 aime24 88.8 %; step 80 aime24 98.8 %; step 100 aime24 99.6 %; step 100 aime25 100.0 %; step 100 math500 100.0 %.

⚠ Read the high-truncation rows as termination artifacts, not as capability being destroyed. step 60 on aime24 is 88.8 % truncated and scores -4.58 pp; step 80 on aime24 is 98.8 % truncated and scores -2.50 pp; step 100 on aime24 is 99.6 % truncated and scores -2.50 pp; step 100 on aime25 is 100.0 % truncated and scores +1.25 pp; step 100 on math500 is 100.0 % truncated and scores -28.75 pp. A policy that never emits a stop token produces a 31,744-token sample with no final answer in it, and the grader scores that as wrong however much the model knows — B-short's ckpt-10 is the control for exactly this: it is 41–60 % truncated and its accuracy is nonetheless preserved, because its answers still appear before the loop begins. The deeper into the sink a checkpoint is, the more of its samples never reach an answer at all.

Only the (checkpoint × benchmark) cells that were approved for this condition appear above: step 60 on aime24; step 80 on aime24. The other combinations were never launched and are not counted as pending.

Transfer ratio at the terminal checkpoint (student gain ÷ the OpenThinker3 pair's teacher gain, five guards)

benchmarkteacher gainstudent gainnumerator CI excludes 0ratioreason
aime24+51.771 [+38.646, +64.375]-2.500 [-7.917, +2.500]Falsenullguard: student_gain = -2.500 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BOTH gains are positive; a non-positive nu
aime25+42.604 [+28.539, +56.667]+1.250 [-2.500, +4.583]Falsenullguard (exp2b, new): student_gain = 1.250 pp has CI95 [-2.500, 4.583] pp which INCLUDES zero — the numerator is statistically indistinguishable from 'n
math500+50.050 [+46.650, +53.400]-28.750 [-35.000, -22.744]Truenullguard: student_gain = -28.750 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BOTH gains are positive; a non-positive n

H5 verdict: NOT SUPPORTED

  • —no_collapse_through_step_100 — FAILED: clip_ratio first exceeded 0.5 at step 48
  • —gain_on_at_least_one_benchmark — FAILED: no benchmark shows a positive paired gain whose 95 % CI excludes 0

13.8 Cross-condition summary — one row per (condition, benchmark)

12 of 21 cells computed, 9 canceled. Each condition's student gain is divided by the teacher gain OF ITS OWN PAIR (run 1 -> r1distillsft; B/C/D and both redirects -> openthinker3sft). A pending cell means the unit is not on the Hub yet; a canceled cell means the unit can never exist because the run was stopped before it saved a checkpoint. Neither is ever an estimate. For B-short the row is its TERMINAL checkpoint (step 10); the per-step table is in the condition's own section.

conditiontierbenchmarkprotocolteacher gain (pp)student gain (pp)ratiowhy not
run 1 — R1-distill shiftrun 1aime24reduced+24.38 [+15.42, +34.17]-5.00 [-11.25, +0.42]nullguard: student_gain = -5.000 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BO
run 1 — R1-distill shiftrun 1aime25reduced+21.15 [+10.21, +33.33]+1.67 [-1.67, +5.42]nullguard (exp2b, new): student_gain = 1.667 pp has CI95 [-1.667, 5.417] pp which INCLUDES zero — the numerator is
run 1 — R1-distill shiftrun 1math500reduced-100-problem-subset+46.50 [+43.20, +49.85]-8.00 [-13.25, -3.25]nullguard: student_gain = -8.000 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BO
B — OpenThinker3 shiftBaime24—+51.77 [+38.65, +64.38]⛔ cancelednullcollapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would o
B — OpenThinker3 shiftBaime25—+42.60 [+28.54, +56.67]⛔ cancelednullcollapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would o
B — OpenThinker3 shiftBmath500—+50.05 [+46.65, +53.40]⛔ cancelednullcollapsed at steps 10-13; its first checkpoint save (step 20) is POST-collapse, so finishing it (~$50) would o
C — OpenThinker3 shift, KL floor = 2.5Caime24—+51.77 [+38.65, +64.38]⛔ cancelednullcollapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint
C — OpenThinker3 shift, KL floor = 2.5Caime25—+42.60 [+28.54, +56.67]⛔ cancelednullcollapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint
C — OpenThinker3 shift, KL floor = 2.5Cmath500—+50.05 [+46.65, +53.40]⛔ cancelednullcollapsed identically to B with the KL coefficient pinned at its maximum; same post-collapse first checkpoint
D — OpenThinker3 shift into Qwen3-4B (thinking)Daime24—+51.77 [+38.65, +64.38]⛔ cancelednullcollapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short becau
D — OpenThinker3 shift into Qwen3-4B (thinking)Daime25—+42.60 [+28.54, +56.67]⛔ cancelednullcollapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short becau
D — OpenThinker3 shift into Qwen3-4B (thinking)Dmath500—+50.05 [+46.65, +53.40]⛔ cancelednullcollapsed at step 17-20 (clip 0.65 -> 0.98) with no pre-collapse checkpoint; the user declined a D-short becau
B-short — OpenThinker3 shift, 10 steps, save every 2B-shortaime24reduced+51.77 [+38.65, +64.38]-3.33 [-7.50, +0.42]nullguard: student_gain = -3.333 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BO
B-short — OpenThinker3 shift, 10 steps, save every 2B-shortaime25reduced+42.60 [+28.54, +56.67]+3.75 [-0.83, +8.75]nullguard (exp2b, new): student_gain = 3.750 pp has CI95 [-0.833, 8.750] pp which INCLUDES zero — the numerator is
B-short — OpenThinker3 shift, 10 steps, save every 2B-shortmath500reduced-100-problem-subset+50.05 [+46.65, +53.40]+0.50 [-3.75, +4.75]nullguard (exp2b, new): student_gain = 0.500 pp has CI95 [-3.750, 4.750] pp which INCLUDES zero — the numerator is
B-lowlr — OpenThinker3 shift at OPTIM_LR 2e-7, 100 stepsB-lowlraime24reduced+51.77 [+38.65, +64.38]-2.50 [-7.92, +2.50]nullguard: student_gain = -2.500 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BO
B-lowlr — OpenThinker3 shift at OPTIM_LR 2e-7, 100 stepsB-lowlraime25reduced+42.60 [+28.54, +56.67]+1.25 [-2.50, +4.58]nullguard (exp2b, new): student_gain = 1.250 pp has CI95 [-2.500, 4.583] pp which INCLUDES zero — the numerator is
B-lowlr — OpenThinker3 shift at OPTIM_LR 2e-7, 100 stepsB-lowlrmath500reduced-100-problem-subset+50.05 [+46.65, +53.40]-28.75 [-35.00, -22.74]nullguard: student_gain = -28.750 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when B
RAFT — rejection-sampling-SFT shift, 100 stepsRAFTaime24full+0.73 [-1.04, +2.60]-1.56 [-4.17, +0.83]nullguard:teacher_gain= 0.729 pp < 1.0 pp — the ratio is not reported
RAFT — rejection-sampling-SFT shift, 100 stepsRAFTaime25full+0.31 [-0.62, +1.35]-2.19 [-5.31, +0.00]nullguard:teacher_gain= 0.312 pp < 1.0 pp — the ratio is not reported
RAFT — rejection-sampling-SFT shift, 100 stepsRAFTmath500full+11.20 [+8.90, +13.55]-1.65 [-3.45, +0.10]nullguard: student_gain = -1.650 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BO

Reading this table. Each row's ratio divides that condition's student gain by the teacher gain of its own pair — run 1 by r1distill_sft, everything else by openthinker3_sft. Crossing the pairs would be meaningless, and the aggregator refuses to do it.

13.9 Every Direct-OPD run in the study, side by side

One row per Direct-OPD training run in the study. primary_gains are the paired per-problem gains vs that condition's own studentinit, taken at the step named in `primarygains_step` — the terminal evaluated checkpoint where its units have landed, otherwise the last checkpoint that HAS been measured — and are null wherever no unit exists at all. Verdicts are computed from the measured training record and the measured gains, never asserted.

collapse onset = the first step whose response_length/clip_ratio exceeds 0.5; "none" means it never did over the steps recorded

runteacher pairstudentlrKL regimestepslength runawaycollapse onsetpeak clipcheckpoints?hypothesisverdict
run 1Qwen2.5-Math-1.5B → DeepSeek-R1-Distill-Qwen-1.5BQwen2.5-7B-Instruct1e-6adaptive [0.5, 2.5]100 / 10060611.000yesH1NOT SUPPORTED
BQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct1e-6adaptive [0.5, 2.5], eps 0.0125 / 10010121.000noH2UNTESTED
CQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct1e-6CONSTANT 2.5 (ADAPTIVEKLLOSSMINCOEF = MAX = 2.5)24 / 10010121.000noH3REFUTED
DQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen3-4B (thinking)1e-6adaptive [0.5, 2.5], eps 0.0122 / 10017170.992noH4REFUTED
B-shortQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct1e-6adaptive [0.5, 2.5], eps 0.0110 / 1010none0.289yesH2 (pre-collapse form)NOT SUPPORTED
B-lowlrQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct2e-7 (logged DEVIATION from the pinned 1e-6)adaptive [0.5, 2.5], eps 0.01100 / 10042481.000yesH5NOT SUPPORTED
RAFTQwen2.5-1.5B-Instruct → RAFT-SFT of itselfQwen2.5-7B-Instruct1e-6adaptive [0.5, 2.5], eps 0.01100 / 100—none0.008yesH6PARTIAL
runcheckpointsgains at stepAIME 2024 gain (pp)AIME 2025 gain (pp)MATH-500 gain (pp)
run 15 merged checkpoints (20/40/60/80/100) on the Hub; 60/80/100 are post-collapse100-5.00 [-11.25, +0.42] (reduced)+1.67 [-1.67, +5.42] (reduced)-8.00 [-13.25, -3.25] (reduced-100-problem-subset)
BNONE — cancelled at step ~24, before the first save at step 20—⛔ no unit can exist⛔ no unit can exist⛔ no unit can exist
CNONE — cancelled at step ~24, before the first save at step 20—⛔ no unit can exist⛔ no unit can exist⛔ no unit can exist
DNONE — cancelled at step ~24, before the first save at step 20—⛔ no unit can exist⛔ no unit can exist⛔ no unit can exist
B-shortcheckpoints 4, 8, 10 evaluated10-3.33 [-7.50, +0.42] (reduced)+3.75 [-0.83, +8.75] (reduced)+0.50 [-3.75, +4.75] (reduced-100-problem-subset)
B-lowlrcheckpoints 20, 40, 60, 80, 100 evaluated100-2.50 [-7.92, +2.50] (reduced)+1.25 [-2.50, +4.58] (reduced)-28.75 [-35.00, -22.74] (reduced-100-problem-subset)
RAFTcheckpoints 20, 40, 60, 80, 100 evaluated100-1.56 [-4.17, +0.83] (full)-2.19 [-5.31, +0.00] (full)-1.65 [-3.45, +0.10] (full)

13.10 The fifth transfer-ratio guard

Run 1's report flagged a gap: the pre-registration guarded the sign of the numerator and the significance of the denominator, but not the significance of the numerator, so a student gain of +1.667 pp [−1.667, +5.417] on AIME 2025 divided into a tidy, meaningless 0.079. exp2b's pre-registration (E3) closes it: the numerator's 95 % CI must also exclude 0.

The new guard is applied uniformly, to run 1 as well — a table that shows six conditions side by side cannot use two different rules. Nothing published is deleted: when a ratio passes the four original guards and fails only the fifth, both aggregate_evals.py and this pipeline still record value_under_exp2a_guards, so run 1's AIME-2025 ratio of 0.079 remains readable and clearly labelled as the number the older rule produced.

Switch state for this build: APPLY_NUMERATOR_CI_GUARD_TO_RUN1 = True.


14. The termination sink

Run 1 found it once and called it a curiosity. exp2b ran the same method with a different teacher pair, with the KL brake pinned at its maximum, and into a different, thinking student — and found it three more times. This section states the mechanism, says what did transfer, and separates the two.

14.1 The mechanism, in four measured steps

(1) The shift reward is negative on the student's own native outputs — for every SFT pair we tried. Direct-OPD's reward is the token-level log-ratio log πpost − log πpre scored on the student's rollouts. At step 1, before any update, that number is what the pair thinks of the student as it already is:

runteacher pairstudentweighted shift reward @ step 1signsteps with a negative mean reward
run 1Qwen2.5-Math-1.5B → DeepSeek-R1-Distill-Qwen-1.5BQwen2.5-7B-Instruct-0.00783negative90 / 100
BQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct-0.03543negative25 / 25
CQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct-0.03543negative24 / 24
DQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen3-4B (thinking)-0.01427negative19 / 22
B-shortQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct-0.03543negative10 / 10
B-lowlrQwen2.5-1.5B-Instruct → OpenThinker3-1.5BQwen2.5-7B-Instruct-0.03543negative100 / 100
RAFTQwen2.5-1.5B-Instruct → RAFT-SFT of itselfQwen2.5-7B-Instruct-0.00479negative100 / 100

Read literally, the objective's only instruction in every one of these runs was stop writing like yourself. It never pointed at a better answer — only away from the current one. That is not a property of the R1 pair: it survived swapping to an Instruct-lineage pair whose π_post is the same family and post-training lineage as the student, and it survived swapping the student for a thinking model whose distribution the shift was supposed to like (H4's premise, refuted at step 1).

(2) There is a region where the reward is exactly zero, and it is non-termination. A log-ratio of 0 means πpost and πpre assign the same probability. The two teachers disagree about how to write; they agree almost everywhere else, and the cheapest place for a language model to reach that agreement is degenerate repetition — an answer, then the same answer again, forever. Every collapsed run shows the same signature: the mean reward rises towards zero, not towards positive, exactly as the rollouts pin at the length cap.

runweighted reward, step 1 → last recordedresponse length, step 1 → lastactor entropy, first → minclip ratio at the last recorded step
run 1-0.0078 → +0.0026339 → 2,6140.156 → 0.0580.699
B-0.0354 → -0.0037339 → 3,3280.156 → 0.0291.000
C-0.0354 → -0.0043339 → 3,3280.156 → 0.0291.000
D-0.0143 → -0.00111,707 → 4,0890.319 → 0.2190.992
B-short-0.0354 → -0.0085339 → 1,5210.156 → 0.0480.289
B-lowlr-0.0354 → -0.0038339 → 3,3280.156 → 0.0391.000
RAFT-0.0048 → -0.0030339 → 5000.156 → 0.0830.000

The sink is worse at evaluation time than in training, and this is what makes the post-collapse numbers look dramatic. Training capped rollouts at 3,328 tokens (4,096 for D); the evaluation cap is 31,744. A checkpoint that pinned at the training cap therefore runs almost ten times further before it is cut off, and the sample the grader sees has no final answer in it at all:

checkpointAIME 2024 truncationMATH-500 truncationMATH-500 paired gain (pp)
B-short ckpt-1044.2 %60.5 %+0.50 [-3.75, +4.75]
run 1 ckpt-10092.1 %96.5 %-8.00 [-13.25, -3.25]
B-lowlr ckpt-10099.6 %100.0 %-28.75 [-35.00, -22.74]

B-short's ckpt-10 is the control that settles the reading. It is the onset step, not a deep sink: 41–60 % truncated, mean output 15–22k tokens — and its accuracy is preserved (+0.50 pp on MATH-500), because its answers still appear before the repetition loop starts. The further into the sink a checkpoint sits, the larger the share of samples in which no answer is ever emitted, and the more negative the score. Those numbers measure termination, not mathematics — which is why this report's verdict rests on the pre-collapse checkpoints, where the model still stops and the gains are simply null.

(3) The KL brake at its maximum does not prevent it. Condition C is condition B with ADAPTIVE_KL_LOSS_MIN_COEF = MAX = 2.5, i.e. the anchor pinned at its strongest value for the whole run.

C held kl_coef at 2.500 for all 24 recorded steps (constant = True) and still crossed the collapse threshold at step 12 — the same step as B (12). H3 REFUTED.

Run 1's adaptive controller did loosen under sustained negative reward (2.5 → 1.012), and that looked like a plausible cause. It is not the cause. The sink is reachable with the leash at its shortest.

(4) A thinking student delays the sink; it does not avoid it. Condition D put the same shift into Qwen3-4B with thinking mode on and a 4,096-token response cap.

B collapsed at step 12, C at step 12, D at step 17 — a delay of 5 steps, on rollouts that started 1,707 tokens long instead of 339. The longer native output buys time; it does not change the destination. And D's weighted reward turned positive only in the last two recorded steps — as it saturated at the cap. The sink is where the teachers agree, and D found it too.

(5) And the learning rate? B-lowlr re-runs B at OPTIM_LR 2e-7 instead of 1e-6 (a logged deviation) on the theory that the grad-norm spikes — 27 at step 5, 13 at step 9, 21 at step 10 in B — are what launch the policy into the sink.

B-lowlr collapsed anyway, at step 48 — 4.0× later than B (step 12), but not never. Lowering the learning rate 5× postponed the sink; it did not avoid it. The mean rollout length crossed twice its step-1 value at step 42, and the weighted reward was rising towards zero (-0.0354 → -0.0038) on exactly the trajectory B, C and D took.

And the spike theory the redirect was built on does not survive its own test. B and C were preceded by grad-norm spikes (27 at step 5, 13 at step 9), which is why lowering the learning rate looked promising — but D collapsed with no spike at all (peak grad norm 4.4), and B-lowlr survived the largest spike in the study (49.8 at step 21) for tens of steps before going anyway. The spikes are a symptom of the same pressure, not the cause of the sink.

14.2 What did transfer: the register, the format, and the length

The failure is specific. It is not "nothing happened" — a great deal happened, very fast, and none of it was capability.

Run 1's student adopted the R1 reasoning voice within twenty steps and abandoned \boxed{} almost entirely in favour of the prompt-instructed Answer: line — away from both of its teachers, which lean \boxed{} on AIME. §8.2 has run 1's table. The same measurement under the OpenThinker pair, for both redirects, is:

benchmarkunitpass`Answer:` line`\boxed{}`nonemean out tok
aime24student_initsample3258.3 %37.5 %4.2 %1,517
aime24pre_teacher_otsample3259.6 %27.2 %13.2 %2,991
aime24post_teacher_otsample322.0 %80.5 %17.5 %18,275
aime24B-short opd_student_otshift-short-ckpt4sample3239.0 %56.4 %4.7 %1,635
aime24B-short opd_student_otshift-short-ckpt8sample3275.2 %19.9 %4.9 %2,014
aime24B-short opd_student_otshift-short-ckpt10sample3291.2 %3.3 %5.4 %16,569
aime24B-lowlr opd_student_otshift_lowlr-ckpt20sample3238.3 %57.5 %4.2 %1,622
aime24B-lowlr opd_student_otshift_lowlr-ckpt40sample3278.2 %17.3 %4.5 %2,088
aime24B-lowlr opd_student_otshift_lowlr-ckpt60sample3291.2 %2.5 %6.2 %29,734
aime24B-lowlr opd_student_otshift_lowlr-ckpt80sample3245.0 %36.7 %18.3 %31,414
aime24B-lowlr opd_student_otshift_lowlrsample320.8 %65.0 %34.2 %31,647
math500student_initsample483.7 %14.6 %1.8 %564
math500pre_teacher_otsample471.5 %19.1 %9.3 %716
math500post_teacher_otsample412.6 %84.2 %3.3 %5,762
math500B-short opd_student_otshift-short-ckpt4sample478.8 %20.1 %1.1 %659
math500B-short opd_student_otshift-short-ckpt8sample498.9 %0.7 %0.4 %870
math500B-short opd_student_otshift-short-ckpt10_m500sub100sample498.0 %1.2 %0.8 %22,020
math500B-lowlr opd_student_otshift_lowlr-ckpt20sample474.0 %23.5 %2.5 %707
math500B-lowlr opd_student_otshift_lowlr-ckpt40sample498.7 %0.9 %0.4 %890
math500B-lowlr opd_student_otshift_lowlr_m500sub100sample40.8 %61.3 %38.0 %31,744

Note what this pair's πpost does on AIME 2024: it is overwhelmingly `\boxed{}` (80.5 % of samples) where the initial student is only 37.5 %. **If Direct-OPD were teaching the student to imitate πpost, \boxed{} is the surface form it would move towards** — and run 1's student famously moved the other way, extinguishing \boxed{} entirely (§8.2).

  • —B-short, terminating checkpoints — \boxed{} share on AIME 2024, initial 37.5 % → step 4: 56.4 % → step 8: 19.9 %; mean output length 1,517 → 1,635 → 2,014 tokens.
  • —B-lowlr, terminating checkpoints — \boxed{} share on AIME 2024, initial 37.5 % → step 20: 57.5 % → step 40: 17.3 %; mean output length 1,517 → 1,622 → 2,088 tokens.
  • —B-short, post-collapse checkpoints (shown for completeness, not interpretable as register): step 10 3.3 % \boxed{} at 44.2 % truncation and 16,569 mean tokens — at that truncation the extractor is finding a \boxed{} somewhere inside a sample that never ends, which says nothing about the policy's answer format.
  • —B-lowlr, post-collapse checkpoints (shown for completeness, not interpretable as register): step 60 2.5 % \boxed{} at 88.8 % truncation and 29,734 mean tokens; step 80 36.7 % \boxed{} at 98.8 % truncation and 31,414 mean tokens; step 100 65.0 % \boxed{} at 99.6 % truncation and 31,647 mean tokens — at that truncation the extractor is finding a \boxed{} somewhere inside a sample that never ends, which says nothing about the policy's answer format.

The move is real, teacher-directed, and non-monotonic. In B-short it peaks at step 4 (56.4 %, +18.9 pp vs the initial student, towards πpost) and in B-lowlr it peaks at step 20 (57.5 %, +20.0 pp vs the initial student, towards πpost) — and then falls back below where it started as the rollouts lengthen. The two redirects trace the same path at different speeds: the lr-2e-7 run reaches the same style waypoints at roughly five times the step count of the lr-1e-6 run. The learning rate rescales the clock, not the path.

Output length, by contrast, is monotonic in every run: it only ever grows, from the first evaluated checkpoint onward, and it is the variable that ends in the sink. Format oscillates; length does not. Whatever the token-level log-ratio is rewarding, it is not "emit π_post's surface form" — and it is certainly not "be more accurate".

14.3 What this says about distilling an SFT policy change into a larger model

The question this campaign was built to answer is whether an established SFT policy change can be transported into a larger student by scoring the student's own tokens with the log-ratio of the two SFT endpoints. Four runs, two teacher pairs, two student families, one KL ablation, and the answer so far is: not as specified, and the obstacle is structural rather than incidental.

The structural problem is that a log-ratio reward is a difference of two models that are not the student. Its sign on the student's own outputs is an empirical fact nobody chooses, and in all four runs it was negative — the pair prefers π_pre on what the student natively writes. A negative-mean reward with an unbounded action space has a trivially reachable optimum: go where the two models agree. For language models that region is degenerate, and the specific degeneracy is never emitting a stop token, because termination is precisely where a chat model and its SFT descendant differ most sharply in probability.

Three claims this study can now exclude as explanations:

  1. 1."The shift had nothing to give." Excluded twice. Run 1's pair carries +24.38 pp on AIME 2024, and the OpenThinker pair carries +51.77 pp [+38.65, +64.38] on the same benchmark at the same token cap for both ends — no budget confound at all.
  2. 2."The adaptive KL controller loosened the leash." Excluded by condition C: the coefficient was constant at its maximum and the run collapsed identically.
  3. 3."A non-thinking student cannot represent a thinking teacher's policy." Excluded by condition D: the thinking student H4 predicted the shift would reward collapsed too — 5 steps after B, not never.

What is left is the objective itself. Concretely, for a rerun to be informative:

changewhat it fixeswhy this study points at it
A length-aware reward, or an explicit termination bonusprices the sinkthe escape region is degenerate non-termination, and nothing in a token-level log-ratio prices length or repetition. 5 of the 7 runs (run 1, B, C, D, B-lowlr) ended with 90 %+ of rollouts pinned at the response cap.
An on-policy KL anchored to the student's OWN init (not a coefficient schedule)keeps the policy inside the region where the reward is even meaningfulC proves the magnitude of the KL term is not the lever; what is missing is an anchor that tracks the student's own initial distribution rather than a scalar penalty.
A shift whose mean is positive on the student BY CONSTRUCTION — e.g. πpre and πpost both derived from the student itself (same-model pre/post SFT)removes the "stop writing like yourself" instruction at the rootthe step-1 sign table in §14.1 is negative for every pair tried, including one deliberately chosen for lineage proximity. Proximity of family was not enough; identity of model is the untested version.
Early stopping on the clip ratio (the collapse-watch rule, as a kill criterion rather than a diagnostic)stops paying for post-collapse stepsevery collapse in this study was visible in response_length/clip_ratio within 1–2 steps of onset, and B/C/D each burned ~10 further steps after it.
A checkpoint save cadence tied to the collapse watch, not to a step gridguarantees a pre-collapse policy exists to evaluateB and C died with no evaluable checkpoint at all because save_freq was 20 and their onset was step 12 / 12. B-short exists only to repair that.

The last row is the cheapest and, in hindsight, the most valuable: the difference between conditions B/C (no data) and B-short (five saved policies for $5.54) is one configuration line.

14.4 Cost of the extension

$332.36 cumulative against the $500 cap (raised from $200 on 2026-08-26), leaving $167.64. Of that, $92.87 is run 1 (phases P1–P6), $192.02 is the exp2b extension (phases P2b, P3b, P4b, P5b) and $47.47 is exp3 (phase P7).

phasewhat it boughtcostshare of total
P1harness + driver builds (CPU, $0)$0.000.0 %
P2run-1 baseline evaluation wave$7.322.2 %
P2bOpenThinker-pair probes$16.465.0 %
P3run-1 OPD smoke$1.070.3 %
P3bexp2b OPD smokes (B, C, D)$3.611.1 %
P4run-1 100-step training$44.4313.4 %
P4bexp2b training: B, C, D (all CANCELED) + B-short + B-lowlr$100.9030.4 %
P5run-1 probes + reduced-protocol endpoints$26.147.9 %
P5bexp2b evaluation wave$71.0521.4 %
P6run-1 checkpoint curve$13.914.2 %
P7exp3: shift anatomy + RAFT pair construction + G2 gate + the RAFT Direct-OPD run + the H6c eval wave$47.4714.3 %
TOTAL$332.36100 %

$46.46 went to jobs that were cancelled (5 of them) — and it matters what that bought. $42.44 is the 3 cancelled training runs (B, C, D). They were stopped once their result was in: their 22–25 recorded steps are the whole of §13.5, and they are what refutes H3 (constant KL) and H4 (thinking student) — two of this campaign's main findings, bought for the price of a fifth of a full run each. Their evidence is training-only because they never reached their first checkpoint save, not because the spend was wasted. The remaining $4.02 — cancelled evaluation jobs: an under-timed attempt that was relaunched successfully, and one mis-shaped job killed on the spot — is the only part that bought nothing. CANCELED jobs are booked on WALL-CLOCK occupancy x flavor rate, not on the platform's running_secs, which it reports as 0 for a cancelled job regardless of the GPU time actually consumed. Each such row says so in its own notes field.

cancelled jobphaseflavorwall-clock hoursbooked
a100x4 full openthinker (B) — CANCELED by user decision at step ~24P4ba100x41.361$13.61
a100x4 full openthinker_klfloor (C) — CANCELED by user decision at step ~24P4ba100x41.361$13.61
a100x4 full openthinkerqwen34b (D) — CANCELED by user decision at step ~24P4ba100x41.522$15.22
exp2b-p5b-ckpt10-m500-sub100 (opdstudentotshift-short-ckpt10_m500sub100, REDUCED protocol, ATTEMPT 1) — CANCELED, undersized timeoutP5ba100-large1.576$3.94
p5b-lowlr-curve-a24-ckpt6080 (combined 2-subfolder job, CANCELED immediately)P5ba100-large0.032$0.08

This is deliberately the larger of the two available bases. The platform reports running_secs = 0 for a cancelled job, which would book every one of them at $0 and understate the extension by $46.46.


15. Answer to the question

Can a policy change learned by SFT be distilled into a larger model with Direct-OPD?

On the evidence of this campaign: no — not as the method is specified, and not because the teachers had nothing to teach. What crossed from the teacher pair to the student was surface form — answer format, register, output length. Held-out mathematical capability did not cross at any checkpoint of any run.

The answer rests on three things, all measured:

  1. 1.The premise was established twice. The paired teacher gain is +24.375 [+15.417, +34.167] p<0.0001 on AIME 2024 for the R1-distill pair and +51.771 [+38.646, +64.375] p<0.0001 for the OpenThinker3 pair, the second with both models at the same token budget. Whatever failed, it was not that the shift encoded nothing.
  2. 2.Nothing improved, anywhere. Across 43 paired held-out measurements — every evaluated checkpoint of every condition, always against the same initial student — 0 are positive with a 95 % CI excluding zero, and 6 are negative with a CI excluding zero. Of the 6 regressions, 4 are post-collapse checkpoints and 2 are small pre-collapse ones. The largest gain measured anywhere is +3.75 pp (B-short step 10, aime25), whose CI includes zero; the largest regression is -28.75 pp (B-lowlr step 100, math500) — and that one is a termination artifact: at 100 % truncation the sample has no final answer for the grader to read. It measures the sink, not a loss of mathematics.
  3. 3.The same failure appeared every time. 5 of the 7 runs (run 1, B, C, D, B-lowlr) crossed the pre-registered collapse watch (clip_ratio > 0.5), and B-short ended at the length-runaway onset before the clip rule could fire in so short a run — under two different teacher pairs, a pinned KL brake, a thinking student and a 5× lower learning rate (§14).

15.0 …and, since exp3, the answer comes with its mechanism

The three statements above are the outcome. exp3 measured why, and the picture that survives has two axes rather than one — a shift's on-supportness (how much of the student's own text it rewards, and what it thinks of stopping) and its magnitude. They are independent, they govern different things, and every pair this campaign has measured falls into one of three regimes:

regimeexample pairson-support?magnitudewhat happens
off-support, largecorpus SFT — R1, OpenThinker3no (23 %/34 % of tokens rewarded; EOS -90)large, ≈ 2.7 /tokentermination collapse; style transfers, capability never does (§14, §17)
on-support, tinyRAFT — π_pre SFT'd on its own verified samplesyes (84.7 % of tokens; EOS -0.006)tiny, 0.099 /tokensafe but null — 100 clean steps, no collapse, no detectable gain (§18)
on-support, largeRL — the pilot's KL-regularized pairyes (58.1 % of tokens)large, RMS 1.02 /tokentransfers ≈ half the teacher gain (the pilot)

So the sharpened answer is: Direct-OPD is not broken and SFT is not disqualified — what fails is a shift that is off the student's support. An SFT pair that is on-support by construction runs safely through the identical channel. It simply had nothing large enough to say. The RL pair gets both properties at once for a structural reason: for KL-regularized RL the log-ratio is the learned advantage divided by β, so on-supportness and magnitude arrive together. Nothing in this campaign shows that an SFT pair cannot occupy that corner — only that neither of the two kinds we could build did.

15.1 What the answer covers

This is an answer about Direct-OPD as pre-registered here, i.e. a token-level reward log πpost − log πpre scored on the student's own rollouts, with a KL penalty, for ≤ 100 steps, using:

  • —two SFT pairs, both 1.5B, both pure-SFT (no RL): a base→reasoning pair (Qwen2.5-Math-1.5B → DeepSeek-R1-Distill-Qwen-1.5B) and an instruct→reasoning pair from the student's own family (Qwen2.5-1.5B-Instruct → OpenThinker3-1.5B);
  • —two students larger than the teachers: Qwen2.5-7B-Instruct (4.7×, non-thinking) and Qwen3-4B (2.7×, thinking);
  • —held-out competition mathematics: AIME 2024 / 2025 (32 samples/problem) and MATH-500 (4 samples/problem), paired per problem, 10,000-resample bootstrap.

Within that box the result is not marginal — it is a null at every checkpoint plus a reproducible, mechanistically-explained training failure.

15.2 What the answer does NOT cover

Six limits, stated so that nobody over-reads the sentence above:

  1. 1.The positive-shift regime is untested. In every run the shift reward's mean on the student's own outputs was negative at step 1 (run 1 -0.0078; B -0.0354; C -0.0354; D -0.0143; B-short -0.0354; B-lowlr -0.0354; RAFT -0.0048). A pair whose log-ratio is positive on the student — which §16 says how to build — is a different experiment, and this campaign says nothing about it.
  2. 2.The clean training window was short. The longest stretch of pre-collapse training any evaluated checkpoint saw is 40 steps at lr 2e-7 (B-lowlr ckpt-40) and 8 steps at lr 1e-6 (B-short ckpt-8). A null over that window bounds what a few steps of this objective can do; it does not bound what a non-collapsing version of it might do over 100.
  3. 3.Both teacher gaps carry an asterisk. Run 1's πpre is context-limited, and the cap-matched sensitivity view (§8.4) shrinks its AIME 2024 gap to a CI that includes zero; the OpenThinker πpost truncates on 19.0 % of AIME 2024 samples even at the full 31,744-token cap, so its gains are lower bounds — the true teacher gap is larger, which only sharpens the contrast with the student's null. Note also how that was discovered: the 8-problem pre-flight probe put π_post's AIME truncation at ≈ 6 %, three to four times below what the full wave measured. Probes measure central tendency, not tails — the same estimation lesson run 1 recorded, now with a second instance. The MATH-500 gap survives every one of these objections intact.
  4. 4.Three conditions were answered only in their training form. B, C and D were cancelled before their first checkpoint save, so H2 and H4 are settled by training evidence (a collapse, and a step-1 reward sign) and not by any evaluation of their policies. B-short exists precisely to supply the pre-collapse policy B never saved.
  5. 5.Post-collapse endpoints use reduced protocols. Non-terminating checkpoints cost 5–10× a normal unit at the 31,744-token cap, so run 1's ckpt-100 and B-short's ckpt-10 were measured at 8 samples/problem on AIME and on a 100-problem MATH-500 subset. Same estimand, wider CIs, always labelled reduced.
  6. 6.One seed per condition. Every run is seed 42 (plus the pilot's seed patch). B-short reproducing B's trajectory step-for-step shows the trajectory is deterministic given the seed; it does not establish that the collapse step is seed-independent.

16. Recommendations for a next attempt

These follow from §14's mechanism and §15's limits. The first is the scientific fix; the rest are what make the experiment survivable and cheap enough to read.

1. Make the shift's mean POSITIVE on the student, by construction. The single fact that explains every run in this study is that log πpost − log πpre was negative on the student's own outputs before a single update. The way to guarantee otherwise is to stop borrowing the pair from a different model: take the student itself as πpre and an SFT of *that student* as πpost (same-model pre/post). The shift then describes a change the student's own distribution can express, and the reward's zero-set is no longer 'wherever these two strangers happen to agree'. Everything else in this list is a guard rail; this is the experiment.

2. Price termination explicitly — a length-aware reward or an EOS bonus. A token-level log-ratio contains no term for stopping, and the cheapest way to drive it to zero is to never stop. 5 of the 7 runs (run 1, B, C, D, B-lowlr) ended with 90 %+ of rollouts pinned at the response cap. A per-sequence bonus on emitting EOS, or a penalty on the clipped fraction, converts the sink from an optimum into a cost. This is cheap to add and directly targets the observed failure.

3. Anchor the KL to the student's own init, not to a coefficient schedule. Condition C settles that the magnitude of the KL term is not the lever: pinned at its maximum 2.5 for every step, it collapsed on the same step as the adaptive run. What is missing is not a bigger penalty but a reference the penalty is measured against — the student's initial policy — so that 'stop writing like yourself' is bounded rather than unbounded.

4. Make the collapse watch a KILL criterion, not a diagnostic. The pre-registered rule (response_length/clip_ratio > 0.5) fired within 1–2 steps of every collapse in this study, and then the runs kept paying: run 1 ran 39 further steps after onset, B ran 13 further steps after onset, C ran 12 further steps after onset, D ran 5 further steps after onset, B-lowlr ran 52 further steps after onset. Stopping at onset would have cost nothing scientifically — the post-collapse checkpoints are the least informative artefacts in the campaign and the most expensive to evaluate (non-terminating checkpoints re-priced the full protocol at $177–208 for a single endpoint). With one caveat this campaign learned the hard way: the driver merges and uploads checkpoints only after training, so B-lowlr was deliberately allowed to finish — cancelling it would have destroyed its two clean pre-collapse policies. An early-stop rule is only safe once item 5 is in place.

5. Tie the checkpoint cadence to the watch, not to a step grid. B and C died with no evaluable policy at all because save_freq was 20 and their onset was step 12. B-short bought five saved policies for $5.54 — one configuration line, and it is the difference between a condition with data and a condition without.

6. Gate on the shift's magnitude and sign BEFORE booking a training run. Every collapse in this campaign was predictable from numbers that cost minutes: the step-1 delta_opd/weighted_reward_mean on the student's own rollouts (negative in all 7 runs), and the unweighted log_ratio_mean (+1.79 for the R1-distill pair, -0.15 for the OpenThinker pair — a 12x smaller signal that collapsed sooner, so magnitude alone does not predict the delay either). A pre-flight gate should require the weighted mean to be positive on a sample of the student's own generations and abort otherwise. The 2-step smoke this campaign already runs measures exactly that number; it was recorded as a note instead of used as a gate.

7. Iterate or scale the RAFT pair — the one direction this campaign has already shown to be safe. exp3 settles item 1 in its safety half: a rejection-sampling SFT pair, on-support by construction, ran the identical channel for 100 steps with a peak clip ratio of 0.0078 and no collapse (§18.3). What it did not settle is transfer, because the shift it produced was tiny — 0.099 |Δ log p| per token, a tenth of the RL pair's RMS — and §18.5 shows the resulting null is exactly the size the pilot's gain-per-shift predicts. One round of rejection-sampling SFT is literally one step of an RL loop, so the natural next move is to keep the construction and grow the magnitude: more rounds (each one re-sampling from the improved model), a larger k per prompt, and a prompt mix that is not skewed easy by the filter (§18.1). That interpolates from the RAFT corner of §15.0's table toward the RL corner, and it is the cheapest available test of whether on-supportness survives iteration — which is the one thing the two-axis picture does not yet tell us.

#changewhat it fixesevidence in this report
1positive-mean shift by construction (same-model pre/post SFT)removes the stop writing like yourself instruction at the root§14.1(1), §15.2(1)
2length-aware reward / EOS bonusprices the sink§14.1(2), §6, F8
3KL anchored to the student's own initbounds the escape§13.5 (condition C), §14.1(3)
4early stop on clip ratio > 0.5stops paying for post-collapse steps§14.1, §11
5save cadence tied to the watchguarantees an evaluable pre-collapse policy§13.5 vs §13.6
6shift-magnitude/sign gate before launchnever books a run whose objective points the wrong way§14.1(1), the smoke metrics in §13.5
7iterated / scaled RAFT — more rounds, more samples per prompt, a harder prompt mixgrows the magnitude while keeping the support§18.5, §18.6

For scale: this campaign cost $332.36, of which $145.33 was training and $46.46 went to cancelled jobs — nearly all of it ($42.44) the B/C/D runs that were stopped once they had refuted H3 and H4, so it is a saving rather than a loss. What items 4 and 5 would have bought is not that money back but the post-onset steps inside every run that ran to completion, plus an evaluable checkpoint for B and C; item 6 is the only one that could have prevented a run from being booked at all.


17. The mechanism, measured

§14 argued from six training runs that the token-level log-ratio is the problem. exp3 Part A measures it directly, without training anything. The same frozen student rollouts are scored under five different teacher pairs, so the only thing that varies across the columns below is which pair you subtract — the text, the tokenizer positions and the correctness verdicts are identical.

17.1 Method

  • —Corpus: the existing student_init generations of Qwen2.5-7B-Instruct on AIME 2024 and MATH-500 — 2,960 frozen rollouts, 2,067,933 scored response positions (2,065,012 ordinary content tokens plus 2,921 special positions, which are excluded from the content statistics and analysed separately as the EOS column).
  • —Models: six 1.5B teachers plus the 7B student, scored in fp32, each reading the student's verbatim token IDs — no re-tokenization, no re-rendering. Every pair is therefore a difference of two log-probabilities over the same index.
  • —What is computed per pair: the per-token log-ratio log πpost − log πpre and its distribution; the ratio at the terminal `<|im_end|>` of rollouts that actually finished (2,921 of them) and at 25,553 sampled interior boundary positions as a control; and the AUROC of the sequence-mean log-ratio against the rollout's own correctness verdict, raw and length-controlled.
  • —Bootstrap: 2000 resamples over problems, seed 42. Only 39 sequences were cap-limited and 17 were truncated, so the corpus is essentially all complete answers.

Everything below is read from `anatomy/stats.json`, which is the canonical artifact: the analysis run's own prose tables pool per-benchmark and differ from it by up to 0.07.

17.2 The five pairs

pairkindtoken meanstudent-weighted mean% tokens > 0terminal EOSinterior EOSAUROC rawAUROC length-controlled
R1corpus SFT (base->chat)-0.111-0.094422.9%-4.238+4.210.6670.747 [0.697, 0.791]
OTcorpus SFT (chat->chat)-0.233-0.193334.3%-90.200-58.670.4740.751 [0.703, 0.792]
RLRL (pilot's working pair)-0.197-0.160458.1%-6.338-4.940.7540.812 [0.777, 0.843]
pilotSFTsmall in-family SFT+0.014+0.012254.7%+0.700-0.120.5470.378 [0.339, 0.424]
RAFTrejection-sampling SFT-0.033-0.008584.7%-0.006-8.470.7430.755 [0.716, 0.790]

Read the two bold columns together. % tokens > 0 is how much of what the student already writes the pair rewards; terminal EOS is what the pair thinks of the student stopping. The corpus pairs reward a quarter to a third of the student's own tokens and punish stopping; the RL pair rewards more than half; the RAFT pair rewards 85 % and is neutral about stopping.

17.3 Where the stop token went

The EOS column above is a ratio. The reason it is so extreme for the corpus pairs is visible only per model — this is mean log p(<|im_end|>) at the genuine end of a completed answer, i.e. how strongly each model believes the answer is over:

modellineagemean log p at the realized token**log p(`<im_end>`) at the true end**at interior positionstreats 151645 as its stop token
student_qwen25_7b_instructqwen-0.081-0.01-36.1yes
qwen25_1_5b_instructqwen-0.176-0.11-34.1yes
raft_qwen25_1_5bqwen-0.208-0.11-42.6yes
qwen_math_1_5bqwen-0.190-18.39-31.8no
pilotsft_deepmath100deepseek-0.292-21.93-27.8no
r1_distill_1_5bdeepseek-0.307-22.63-27.6no
justrl_deepseek_1_5bdeepseek-0.513-28.97-32.6no
openthinker3_1_5bqwen-0.536-90.31-92.8no

OpenThinker3 is the extreme case: -90.3, and stats.json records that not one of its terminal positions has a positive log-ratio. A model trained on long chain-of-thought corpora has effectively unlearned the stop token on this kind of text. Subtract its pre-teacher — which stops perfectly happily (-0.109) — and the reward at every EOS position is a cliff. That is condition B's step-10 collapse, measured on frozen text before any training happened.

Lineage caveat, and it matters for exactly one row. Token id 151645 is <|im_end|> to a Qwen-lineage model and <|Assistant|> to a DeepSeek-lineage one, so an EOS ratio is only well posed when both members of the pair read it the same way:

pairπ_pre lineageπ_post lineageEOS ratio well posed?
R1qwendeepseekno — mixed lineage
OTqwenqwenyes
RLdeepseekdeepseekyes
pilotSFTdeepseekdeepseekyes
RAFTqwenqwenyes

R1's EOS numbers are therefore not a clean ratio (Qwen2.5-Math → DeepSeek-R1-distill crosses the lineage boundary), and the report does not lean on them; the per-model column above, which is a single model's own log-probability, is well posed for every row. The RAFT pair is recorded in stats.json as mixed-lineage too, but that is a bug in the anatomy run's lineage map, not a property of the pair: the RAFT teacher is an SFT of Qwen2.5-1.5B-Instruct whose tokenizer files are sha256-identical to it, so both members are Qwen and its EOS ratio is well posed. Fixed in code/anatomy/analyze_anatomy.py; the published stats.json is left as the original artifact and corrected here.

17.4 The pre-registered predictions, and the four that failed

§F2 pre-registered five predictions before any of this was measured. They are reprinted exactly as the anatomy run recorded them — four of the six coded checks FAILED, and the failures are the most useful part of this section.

checkverdictrulemeasured
P1_corpus_EOS_much_less_than_0_at_terminalPASSevery corpus pair mean <= -0.5R1 -4.238; OT -90.200
P1_corpus_EOS_much_less_than_0_at_interior_boundariesFAILevery corpus pair mean <= -0.5R1 +4.044; OT -63.375
P2_corpus_AUROC_approx_0.5_after_length_controlFAILabs(AUROC-0.5) < 0.05 or CI95 contains 0.5R1 0.747; OT 0.751
P3_RL_EOS_approx_0_or_positive_at_terminalFAILabs(mean) < 0.5 or mean > 0RL -6.338
P3_RL_EOS_approx_0_or_positive_at_interior_boundariesFAILabs(mean) < 0.5 or mean > 0RL -3.859
P4_RL_least_negative_student_weighted_meanFAILRL has the maximum (least negative) student-weighted mean of all pairsR1 -0.094; OT -0.193; RL -0.160; pilotSFT +0.012; RAFT -0.009
P5_RAFT_patterns_with_RLPASSrequires the Part-B RAFT teacher—

What each failure taught:

  1. 1.P1-interior FAILED for R1 — its interior-boundary ratio is positive where the terminal one is sharply negative. That is the mixed-lineage artifact above: at interior positions the two models are being asked about different tokens. Flagged, not repaired.
  2. 2.P2 FAILED, and this is the important one. The prediction was that a corpus pair's sequence-mean log-ratio would be uninformative about correctness once length is controlled. It is not: after length-decile pooling the corpus pairs reach R1 0.747; OT 0.751, against the RL pair's 0.812. The corpus shifts DO carry a sequence-level correctness signal — the pilot's "it is only a length artifact" reading does not replicate here. What they lack is a way for a token-level optimizer to reach it: the anti-EOS and negative-mass gradient dominates every update long before any sequence-level signal could be exploited. A sequence-level objective on the same quantity is a different, untried experiment.
  3. 3.P3 FAILED — the RL pair is anti-EOS too (terminal -6.34, worse than R1's -4.24) and it did not collapse in the pilot. So an anti-EOS shift is not sufficient for the sink. Combined with §14, the honest statement is that anti-EOS magnitude orders collapse speed (OT −90 → step 10–13; R1 −4.2 → step 61) without being the whole cause.
  4. 4.P4 FAILED — the RL pair does not have the least-negative student-weighted mean. R1 -0.0944; OT -0.1933; RL -0.1604; pilotSFT +0.0122; RAFT -0.0085. The only positive-mean pair is pilotSFT — which is also the only pair in the whole campaign that transferred anything in-distribution (+2.59 pp in the pilot). Mean sign alone is the wrong summary statistic; what separates the working pair from the corpus pairs is the positive-token fraction (RL 58.1% vs corpus 22.9%/34.3%) together with EOS neutrality.
  5. 5.P5 (the RAFT pair) — the coded check cannot fail. The anatomy script records P5_RAFT_patterns_with_RL: PASS, but its rule is literally "requires the Part-B RAFT teacher" — it is a presence check, and it would have read PASS whatever the numbers said. This report therefore evaluates the substantive criteria itself, from stats.json, against thresholds fixed in code:
P5 criterionthresholdRAFTRL, for referenceverdict
frac_gt0_at_least_half0.50+0.8467+0.5814HELD
terminal_eos_near_zero1.00-0.0060-6.3385HELD
student_weighted_mean_closer_to_zero_than_corpuscloser to 0 than either corpus pair-0.0085R1 -0.0944; OT -0.1933HELD

P5, evaluated on the numbers: HELD. The RAFT pair does pattern with the RL pair on every criterion the restated prediction named — and on two of the three it is milder than RL, not merely similar. §18 is what happened when that pair was put through the same training channel that broke every corpus pair.

[image]

[image]

[image]


18. The RAFT condition — an SFT pair that is on-support by construction

§17 says the corpus pairs fail because of what the density ratio means, not because SFT is the wrong algorithm. §F1 turns that into a falsifiable design: build a πpost by **rejection-sampling SFT** — SFT in mechanics, but πpost ≈ πpre · exp(advantage) in distribution, because the training data is πpre's own verified samples. If the mechanism claim is right, that pair should behave like the RL pair and not like the corpus pairs, through the identical training channel.

18.1 Building the pair

  • —π_pre Qwen/Qwen2.5-1.5B-Instruct @ 989aa798 — the same π_pre as condition B.
  • —π_post cmpatino/Qwen2.5-1.5B-Instruct-DeepMath-RAFT @ d46294c2a827c8558547c8ebac96a49b7a8410bb, an SFT of π_pre on its own filtered samples.
  • —Stage 1 (gate G1) — k=4 completions per prompt from πpre over the pilot's 6,400 `sfttrain prompts (T 0.7, top-p 0.95, cap 2,048), verified with the harness's own grader block (byte-identical copy, pinned by sha256), keeping at most one correct sample per prompt: **25,600 samples → 9,081 correct (sample accuracy 0.355) → 4,187 kept (pass@4 0.654)**, against a pre-registered floor of 2,000. **G1 PASS.** Ground truth was re-derived from DeepMath-103K @ 5cf055d1` and the join was gated on 6,400/6,400 question-text and qhash matches.
  • —The keep rate falls with difficulty — 0.894 at difficulty bin 3.0 down to 0.55–0.62 at bins 7.5–9.5 — so the RAFT training set is easier-skewed relative to sft_train. That is what rejection sampling does; it is stated rather than hidden, and it bounds how much this teacher could ever teach about hard problems.
  • —Stage 2 — assistant-masked SFT, lr 1e-5 cosine (inside the pre-registered [5e-6, 2e-5]), 2 epochs = 124 steps, global batch 64, fp32 master weights. Validation loss 0.16881 → 0.15517 (25 %) → 0.15384 (50 %) → 0.16230 (75 %) → 0.16166 (100 %): the gate passes, but the minimum is at the end of epoch 1 — the second epoch mildly over-fits. 2 epochs was fixed before launch, so the root checkpoint is the pre-registered π_post and no post-hoc selection was done; checkpoint-50pct exists in the repo if anyone ever wants the lower-loss variant, and choosing it would be a new, logged decision.
  • —The shift this produces is small: mean |Δ log p| per token 0.099 — roughly 27× smaller than the corpus pairs' ≈ 2.7. Hold on to that number; §18.5 needs it.

18.2 Gate G2 — does the RAFT teacher actually know more?

benchmarkpassπ_preπ_post^RAFTpaired gain (pp)
math500greedy45.4053.80+8.400 [+4.200, +12.800] p<0.0001
math500sample4 (PRIMARY)38.7549.95+11.200 [+8.900, +13.550] p<0.0001
teacher_evalgreedy37.1153.32+16.211 [+11.523, +20.898] p<0.0001
teacher_evalsample4 (PRIMARY)34.5252.54+18.018 [+15.137, +20.850] p<0.0001
aime24sample32 (PRIMARY)2.192.92+0.729 [-1.042, +2.604] p=0.3582
aime25sample32 (PRIMARY)0.420.73+0.312 [-0.625, +1.354] p=0.4852

G2 PASS — PREREGISTRATION F3 G2: paired MATH-500 gain of RAFT - pipre > 0 with a 95 % CI excluding 0. MATH-500 `sample4` **+11.200** [+8.900, +13.550] p<0.0001. The in-distribution probe (`teachereval`, the pilot pool's held-out 512-prompt split, disjoint from the RAFT training prompts) moves further still. Part of the MATH-500 gain is format compliance — extraction success rose 0.918 → 0.994 — but conditional-on-complete accuracy also rose (39.10 → 50.94), so it is not only format. The RAFT teacher is a genuinely better model than its own pre-teacher, which is the whole point: an on-support shift that nevertheless carries real capability.

A second, sharper check ran alongside it. Measuring the log-probability of the terminal <|im_end|> on 32 held-out completions — byte-identically the ones the training run held out — gives π_pre −0.074, RAFT −0.015, Δ = +0.059 (fp32). The shift at the stop token is not merely small, it is positive: training a model on its own complete samples slightly rewards stopping. The corpus comparator, OpenThinker3, measures −90.3 (§17.3). The anatomy's own frozen-rollout measurement agrees: the RAFT pair's terminal EOS ratio is -0.006.

18.3 Training — H6a and H6b

H6a — the sign of the reward on the student's own outputs. Step-1 delta_opd/weighted_reward_mean = -0.0048. The strict clause (≥ 0) FAILED; the pre-registered weaker clause HELD — it is 7.4× closer to zero than condition B's -0.0354 under the same π_pre, the same student and the same configuration. Verdict: PARTIAL — the shift is far more on-support than any corpus pair, and still not positive.

Student checkpoints: `cmpatino/Qwen2.5-7B-Instruct-DirectOPD-RAFTShift-100` @ `de65f5a389be54f1` (pinned (training COMPLETE 2026-08-31, job 6a958690)); per-step record from `metrics.jsonl (model repo)`.

H6b — no termination collapse: HELD. clipratio stayed at or below 0.5 for all 100 steps (peak 0.0078), and the rollout length never doubled. This is **the first Direct-OPD run in the campaign with an SFT teacher pair that did not collapse.** Same student, same πpre, same 768/3328 lengths, same lr, same adaptive KL, same 100 steps as condition B — which collapsed at step 12. The only difference is where π_post's training data came from.

runteacher pairπ_post's datacollapse onsetpeak clip ratiorollout length, step 1 → end
run 1Qwen2.5-Math-1.5B → DeepSeek-R1-Distill-Qwen-1.5Bsomebody else's corpusstep 611.0000339 → 2,614
BQwen2.5-1.5B-Instruct → OpenThinker3-1.5Bsomebody else's corpusstep 121.0000339 → 3,328
CQwen2.5-1.5B-Instruct → OpenThinker3-1.5Bsomebody else's corpusstep 121.0000339 → 3,328
DQwen2.5-1.5B-Instruct → OpenThinker3-1.5Bsomebody else's corpusstep 170.99221,707 → 4,089
B-shortQwen2.5-1.5B-Instruct → OpenThinker3-1.5Bsomebody else's corpusnone0.2891339 → 1,521
B-lowlrQwen2.5-1.5B-Instruct → OpenThinker3-1.5Bsomebody else's corpusstep 481.0000339 → 3,328
RAFTQwen2.5-1.5B-Instruct → RAFT-SFT of itselfπ_pre's own verified samplesnone0.0078339 → 500

The RAFT run's rollouts never approached the 3,328-token training cap; its peak clip ratio over 100 steps is 0.0078, against 1.000 for every collapsing run. The anti-EOS driver §17 measured is simply absent for this pair.

One metric to read carefully. verl's delta_opd/log_ratio_pos_frac reads ≈ 0.08 on this run's training rollouts, while §17 reports 84.7 % of tokens positive for the same pair. They are different quantities: verl's is computed over the top-16 candidate tokens at each position, the anatomy's over the realized token. Neither is wrong; they must not be compared.

18.4 H6c — did anything transfer?

Because the model terminates, every checkpoint could be measured at the full protocol (MATH-500 500 problems; AIME 2024/2025 32 samples), with the AIME 2024 curve at the pre-registered 8-sample diagnostic protocol.

Note on the protocol column below: for this condition reduced marks only the 8-sample AIME curve points — it is never a truncation-forced fallback, unlike §7 and §13.7. The truncation line under the table is the proof: every unit here truncates on 0.1–2.5 % of samples, so nothing was measured on unfinished answers.

100 of 100 pre-registered steps recorded (source: metrics.jsonl (model repo)). Collapse onset (first clip_ratio > 0.5): none; length-runaway marker (mean rollout length past 2x its step-1 value): none; peak clip ratio 0.008. Grad-norm peak 3.5 at step 1; KL coefficient 2.475 → 0.915; weighted shift reward -0.0048 → -0.0030, negative on 100 of 100 steps.

conditionstepresponse len (mean)clip ratioactor entropyweighted shift rewardKL coefgrad norm
RAFT13390.0000.156-0.00482.4753.52
RAFT53860.0000.087-0.00202.3771.06
RAFT103900.0000.100-0.00242.2611.02
RAFT154610.0000.092-0.00232.1500.70
RAFT204310.0000.097-0.00202.0450.78
RAFT405610.0040.091-0.00201.6720.74
RAFT605690.0000.099-0.00231.3680.67
RAFT805870.0040.138-0.00321.1190.68
RAFT1005000.0000.185-0.00300.9150.64

Held-out gains, per evaluated checkpoint (paired per-problem bootstrap vs student_init; the protocol is detected from each unit, never assumed)

stepunitbenchmarkprotocolpairingckpt (pp)baseline (pp)gain (pp)n
20opd_student_raftshift-ckpt20aime24reducedseeds 0–7 (PRIMARY)11.2513.75-2.500 [-7.917, +2.500] p=0.288430
20opd_student_raftshift-ckpt20aime24reducedvs the full 32-sample baseline (unpaired-in-samples)11.2512.19-0.938 [-5.104, +2.708] p=0.641630
20opd_student_raftshift-ckpt20math500fullgreedy76.0075.40+0.600 [-2.000, +3.200] p=0.6066500
20opd_student_raftshift-ckpt20math500fullsample4 (PRIMARY)73.6574.35-0.700 [-2.300, +0.900] p=0.3660500
40opd_student_raftshift-ckpt40aime24reducedseeds 0–7 (PRIMARY)11.6713.75-2.083 [-6.667, +2.500] p=0.321630
40opd_student_raftshift-ckpt40aime24reducedvs the full 32-sample baseline (unpaired-in-samples)11.6712.19-0.521 [-3.750, +2.708] p=0.707230
40opd_student_raftshift-ckpt40math500fullgreedy76.0075.40+0.600 [-2.400, +3.600] p=0.6422500
40opd_student_raftshift-ckpt40math500fullsample4 (PRIMARY)74.6074.35+0.250 [-1.400, +1.900] p=0.7548500
60opd_student_raftshift-ckpt60aime24reducedseeds 0–7 (PRIMARY)12.0813.75-1.667 [-7.083, +3.750] p=0.501230
60opd_student_raftshift-ckpt60aime24reducedvs the full 32-sample baseline (unpaired-in-samples)12.0812.19-0.104 [-4.271, +3.857] p=0.927630
60opd_student_raftshift-ckpt60math500fullgreedy73.6075.40-1.800 [-4.600, +1.000] p=0.1930500
60opd_student_raftshift-ckpt60math500fullsample4 (PRIMARY)74.0074.35-0.350 [-2.100, +1.300] p=0.6432500
80opd_student_raftshift-ckpt80aime24reducedseeds 0–7 (PRIMARY)11.6713.75-2.083 [-7.083, +2.917] p=0.345630
80opd_student_raftshift-ckpt80aime24reducedvs the full 32-sample baseline (unpaired-in-samples)11.6712.19-0.521 [-3.542, +2.604] p=0.688630
80opd_student_raftshift-ckpt80math500fullgreedy73.8075.40-1.600 [-4.600, +1.400] p=0.2678500
80opd_student_raftshift-ckpt80math500fullsample4 (PRIMARY)73.4074.35-0.950 [-2.750, +0.800] p=0.2870500
100opd_student_raftshiftaime24fullseeds 0–31 (PRIMARY)10.6212.19-1.562 [-4.167, +0.833] p=0.179430
100opd_student_raftshiftaime25fullseeds 0–31 (PRIMARY)4.796.98-2.188 [-5.312, +0.000] p=0.036830
100opd_student_raftshiftmath500fullgreedy74.6075.40-0.800 [-4.000, +2.200] p=0.5734500
100opd_student_raftshiftmath500fullsample4 (PRIMARY)72.7074.35-1.650 [-3.450, +0.100] p=0.0642500

Truncation at the 31,744-token evaluation cap, per checkpoint: step 20 aime24 2.5 %; step 20 math500 0.1 %; step 40 aime24 1.2 %; step 40 math500 0.2 %; step 60 aime24 2.5 %; step 60 math500 0.1 %; step 80 aime24 2.1 %; step 80 math500 0.2 %; step 100 aime24 2.4 %; step 100 aime25 1.0 %; step 100 math500 0.4 %.

Only the (checkpoint × benchmark) cells that were approved for this condition appear above: step 20 on aime24, math500; step 40 on aime24, math500; step 60 on aime24, math500; step 80 on aime24, math500. The other combinations were never launched and are not counted as pending.

Transfer ratio at the terminal checkpoint (student gain ÷ the RAFT pair's teacher gain, five guards)

benchmarkteacher gainstudent gainnumerator CI excludes 0ratioreason
aime24+0.729 [-1.042, +2.604]-1.562 [-4.167, +0.833]Falsenullguard:teacher_gain= 0.729 pp < 1.0 pp — the ratio is not reported
aime25+0.312 [-0.625, +1.354]-2.188 [-5.312, +0.000]Falsenullguard:teacher_gain= 0.312 pp < 1.0 pp — the ratio is not reported
math500+11.200 [+8.900, +13.550]-1.650 [-3.450, +0.100]Falsenullguard: student_gain = -1.650 pp is NOT POSITIVE — the pre-registration reports the transfer ratio only when BOTH gains are positive; a non-positive nu

H6c: FAILED — none of the 11 measured cells shows a positive paired gain whose 95 % CI excludes 0.

Every one of the measured cells is flat or slightly negative, none has a CI excluding 0 on the positive side, and there is no trend with training step: the curve wanders inside ±1 pp from step 20 to step 100. Diagnostics confirm the model is healthy rather than damaged — truncation 0.05–2.5 % everywhere, none-extractor share stable at 1–2 % on MATH-500 — so this is a null, not a repeat of §14's termination artifact.

18.5 What the null means — and why it is the expected size

§F3 pre-registered the failure read: "H6a,b true but H6c false → the channel is safe but this shift is too small to detect." That is exactly the branch we are in, and it can be checked quantitatively rather than accepted as an excuse.

The pilot's RL pair transferred about +8.75 pp per unit RMS shift, at an RMS of 1.02 per token. The RAFT shift measures 0.099 |Δ log p| per token — about a tenth. Scaling the pilot's gain-per-shift gives an expected student gain below 1 pp, while this condition's MATH-500 CI is roughly ±1.8 pp wide. The null is quantitatively consistent with the mechanism, not evidence against it: a shift this small could not have been detected at this sample size even if it transferred at the RL pair's efficiency.

18.6 The two-axis picture

Putting §17 and §18 together, the campaign resolves into two independent axes, and every run in it sits where those axes put it:

pairon-support? (% tokens > 0, EOS)magnitude (Δ log p/token)outcome
corpus SFT (R1, OT)no — 23–34 % of tokens, EOS −4 to −90large, ≈ 2.7collapse; style transfers, capability does not
RAFT (own-samples SFT)yes — 85 % of tokens, EOS ≈ 0tiny, 0.099safe but null; no collapse, no detectable gain
RL (the pilot's pair)yes — 58 % of tokenslarge, RMS 1.02transfers, ≈ half the teacher gain

On-supportness governs stability; magnitude governs transfer. They are separate properties, and this campaign now has a condition at each corner it could reach. The RL pair gets both at once for a structural reason rather than a lucky one: for KL-regularized RL, log πpost − log πpre is the learned advantage divided by β, so the shift is on-support because it came from the policy itself and large because the advantage is what the training optimised. Corpus SFT buys magnitude without support; a single round of RAFT buys support without magnitude.

The open question this leaves is a specific one, not a shrug: iterated or scaled RAFT — more rounds, more samples per prompt, a harder prompt mix — is literally one step of an RL loop repeated, and it is the natural interpolation from the RAFT corner toward the RL corner. Whether the shift magnitude grows fast enough, while the on-supportness survives the iteration, is the next experiment (§16).