dougalldeepmind/2026-08-03-qwen36-nika-sft-tulu-toolcall-80-20-only-closing-think-tokens-loss
nika-sft-tulu-toolcall-80-20-only-closing-think-tokens-loss
LoRA adapter for Qwen/Qwen3.6-27B. SFT on an 80% TULU3 replay / 20% agentic tool-calling difficult-advice mixture, with the loss masked to assistant tokens and a think-loss rule that masks the `<think>` opener and always supervises `</think>`.
The <think> opener is Qwen3.6's marker, injected as a prefill at inference time, so the model should be conditioned on it rather than trained to emit it — training a model to emit one is the documented reasoning-collapse pattern. </think> is the opposite: closing the block is behaviour the model must actually learn, so it always carries loss, whether or not there is reasoning between the tags.
Relationship to the both-think-tokens arm
Direct counterpart to `nika-sft-tulu-toolcall-80-20-both-think-tokens-loss`, which is deprecated. Identical mixture — the same published file, not a rebuild — and identical hyperparameters. The think-loss rule is the only difference, so any behavioural difference between the two is attributable to it alone. The deprecated arm supervised all 2,449 empty <think></think> blocks, training the model to emit the non-thinking marker.
Training data
`LASR-Callum/2026-08-03-tulu-toolcall-80-20-mixture` — 1,791 examples, 1,496,873 tokens, 79.98% TULU3 / 20.02% agentic.
Hyperparameters
The rule, exactly
Implemented as one predicate: mask tokens lying wholly inside a span covering the <think> literal plus one following newline. Qwen's tokenization makes that cover both cases with no special-casing for empty blocks:
empty block: <think> · \n\n LOSS </think> LOSS
(\n\n is ONE token, only partly inside the span, so it keeps loss)
real reasoning: <think> · \n · Let me check. LOSS </think> LOSS
(that \n is its own token, wholly inside the span, so it is masked)Verified against the real tokenizer: <think>=248068, </think>=248069, \n=198, \n\n=271.
Verification before training
scratch/verify_mask.py, run on real mixture rows with a parser that re-derives role regions independently rather than re-running the masking code:
- 120 rows (72 agentic, 48 TULU3), 251,669 tokens, over-weighting multi-turn and tool-call-heavy rows
- `<think>` openers carrying loss: 0. `</think>` carrying loss: 597/597.
- 0 supervised tokens on system/user/tool content; 0 outside any assistant span
- 95 multi-turn rows: the window reaches 572 assistant turns, all 572 supervised
- supervised 68.1% of sampled tokens, 74.0% over the full mixture
The delta against the deprecated arm is exactly accounted for: 1,111,004 → 1,108,331 = 2,673 tokens, being 2,561 <think> openers plus 112 reasoning-turn newlines. Nothing else changed.
Sequence length
4096, not the sibling arms' 2048/3072. At 2048 only 80.4% of the agentic corpus survives and 11 of its <tool_call> spans are severed, silently. At 4096 the corpus is whole: longest row 3,989 tokens, 0 truncated, 0 of its 92 <tool_call> spans severed.
Status
Trained and published. Not yet evaluated — no agentic-misalignment or capability numbers exist for it yet. The obvious first check, given the mixture is 95.6% empty think blocks, is whether it still reasons and still emits well-formed tool calls — and how it compares to the deprecated arm on exactly that.
