dougalldeepmind/2026-08-03-qwen36-nika-sft-tulu-toolcall-80-20-both-think-tokens-loss
nika-sft-tulu-toolcall-80-20-both-think-tokens-loss
DEPRECATED (2026-08-03). This arm trained loss on both<think>and</think>. Because 95.6% of this mixture's think blocks are empty, that trains the model to emit Qwen3.6's non-thinking marker - the documented reasoning-collapse pattern. The rule was removed from core the same day. Superseded by `nika-sft-tulu-toolcall-80-20-only-closing-think-tokens-loss`, which masks the<think>opener and always supervises</think>- same mixture, same hyperparameters, so the two differ only in the think-loss rule. Kept for provenance and as the control in that comparison. Not recommended for use.
LoRA adapter for Qwen/Qwen3.6-27B. SFT on an 80% TULU3 replay / 20% agentic tool-calling difficult-advice mixture, with the loss masked to assistant tokens only and a <think> block on every assistant turn.
This replaces a run retracted on 2026-08-03. The retracted arm had two defects: it trained full-sequence rather than masking the loss to assistant tokens, and it did not guarantee a think block on every assistant turn. Its loss curve fell 2.753 → 1.057 and looked textbook the whole time — a mask defect does not show up in the loss. Both defects are fixed here and the mask was verified against real batches before training started (numbers below).
Training data
`LASR-Callum/2026-08-03-tulu-toolcall-80-20-mixture` — 1,791 examples, 1,496,873 tokens, 79.98% TULU3 / 20.02% agentic.
Hyperparameters
Sequence length: why 4096
At 2048 only 80.4% of the agentic corpus survives and 11 of its <tool_call> spans are severed, inside exactly the long conversations the tool calls live in. Truncation is silent — nothing errors. At 4096 the agentic corpus is whole: longest row 3,989 tokens, 0 rows truncated, 0 of its 92 <tool_call> spans severed.
Label-mask verification
Run before training, on real mixture rows through the real collator, with an independent parser that re-derives each conversation's role regions rather than re-running the masking code (scratch/verify_mask.py):
- 120 rows sampled (72 agentic, 48 TULU3), deliberately over-weighting long multi-turn and tool-call-heavy rows — 251,669 tokens.
- 0 supervised tokens on system, user or tool content, and 0 outside any assistant span.
- 95 multi-turn rows checked; the window reaches 572 assistant turns and all 572 are supervised (not just each row's first turn).
- Supervised fraction 68.3% of tokens on the sample, 74.2% over the full mixture. Assistant content is 69.4% of the same text by character — the ceiling the mask could reach, and what it tracks. High because this corpus is mostly assistant text, not because the mask is inert.
Why this is deprecated: 95.6% of think blocks are empty
Every assistant turn carries a think block and 2,449 of 2,561 are empty — and under assistant-only masking those empty blocks are supervised. That is the intended experiment (an invariant tag structure across the whole corpus), but it is also the documented empty-think collapse by construction. If this adapter stops reasoning, check this first, not last.
Status
Trained and published (public). Not yet evaluated — no agentic-misalignment or capability numbers exist for it yet.
