CoolFace
Modelpublic

dougalldeepmind/2026-08-03-qwen36-nika-sft-tulu-toolcall-80-20-both-think-tokens-loss

sourceHugging Faceotherupdated 27d agoView on Hugging Face
0likes22downloads
Model Card

nika-sft-tulu-toolcall-80-20-both-think-tokens-loss

DEPRECATED (2026-08-03). This arm trained loss on both <think> and </think>. Because 95.6% of this mixture's think blocks are empty, that trains the model to emit Qwen3.6's non-thinking marker - the documented reasoning-collapse pattern. The rule was removed from core the same day. Superseded by `nika-sft-tulu-toolcall-80-20-only-closing-think-tokens-loss`, which masks the <think> opener and always supervises </think> - same mixture, same hyperparameters, so the two differ only in the think-loss rule. Kept for provenance and as the control in that comparison. Not recommended for use.

LoRA adapter for Qwen/Qwen3.6-27B. SFT on an 80% TULU3 replay / 20% agentic tool-calling difficult-advice mixture, with the loss masked to assistant tokens only and a <think> block on every assistant turn.

This replaces a run retracted on 2026-08-03. The retracted arm had two defects: it trained full-sequence rather than masking the loss to assistant tokens, and it did not guarantee a think block on every assistant turn. Its loss curve fell 2.753 → 1.057 and looked textbook the whole time — a mask defect does not show up in the loss. Both defects are fixed here and the mask was verified against real batches before training started (numbers below).

fieldvalue
experimentTool-calling 80/20 SFT arm of the Teaching Claude Why replication: does an invariant <think> tag structure, trained with assistant-only loss, reduce agentic misalignment without breaking tool calling?
date_generated2026-08-03
constitutionconstitutions/claude_approved_constitution.md (in the source repo, commit below) — the agentic 20% of the training mixture derives from it. The TULU3 80% relates to no constitution.
source_repoMatthew-Bozoukov/teaching_claude_why_replication, branch nika-sft-tulu-toolcall-80-20, commit 8307750968f132df00832938ea50fe6b196398b9
modelsBase Qwen/Qwen3.6-27B, loaded bf16 via AutoModelForImageTextToText (it ships as a vision-language checkpoint).
generation_configTraining, not sampling: see hyperparameters below.
schemaStandard PEFT adapter — adapter_model.safetensors + adapter_config.json + tokenizer files.
provenancepython scripts/train_lora.py --config configs/train_lora_toolcall_80_20_thinkall.yaml on 1×H100 80GB. That config no longer runs; it is kept, deprecated, at scratch/deprecated/lora_qwen36_toolcall_80_20_both_think.yaml, and the removed masking code at scratch/deprecated/think_loss_legacy.py.

Training data

`LASR-Callum/2026-08-03-tulu-toolcall-80-20-mixture` — 1,791 examples, 1,496,873 tokens, 79.98% TULU3 / 20.02% agentic.

Hyperparameters

LoRAr=32, α=64, dropout=0.05
target modules`model\.languagemodel\..*\.(qproj\k_proj\v_proj\o_proj\gate_proj\up_proj\down_proj)$`
precisionbf16 (not 4-bit: bitsandbytes does not reliably cover this model's linear-attention layers)
epochs1
batch1 × grad-accum 16
lr1e-4, cosine → 0, warmup 3%
maxseqlen4096 (sibling arms use 2048; see below)
packingoff
loss maskingassistant tokens only
steps112
loss0.853 → 0.762 over 22 points logged every 5 steps. Note the level: the retracted full-sequence run started at 2.753. Masked loss sees only assistant tokens, which the base model already predicts far better than prompt tokens, so the ~3x lower start is itself corroboration that the mask changed which tokens are trained.
wall clock1h23m (112 steps at ~40.6 s/step) on 1×H100 80GB

Sequence length: why 4096

At 2048 only 80.4% of the agentic corpus survives and 11 of its <tool_call> spans are severed, inside exactly the long conversations the tool calls live in. Truncation is silent — nothing errors. At 4096 the agentic corpus is whole: longest row 3,989 tokens, 0 rows truncated, 0 of its 92 <tool_call> spans severed.

Label-mask verification

Run before training, on real mixture rows through the real collator, with an independent parser that re-derives each conversation's role regions rather than re-running the masking code (scratch/verify_mask.py):

  • —120 rows sampled (72 agentic, 48 TULU3), deliberately over-weighting long multi-turn and tool-call-heavy rows — 251,669 tokens.
  • —0 supervised tokens on system, user or tool content, and 0 outside any assistant span.
  • —95 multi-turn rows checked; the window reaches 572 assistant turns and all 572 are supervised (not just each row's first turn).
  • —Supervised fraction 68.3% of tokens on the sample, 74.2% over the full mixture. Assistant content is 69.4% of the same text by character — the ceiling the mask could reach, and what it tracks. High because this corpus is mostly assistant text, not because the mask is inert.

Why this is deprecated: 95.6% of think blocks are empty

Every assistant turn carries a think block and 2,449 of 2,561 are empty — and under assistant-only masking those empty blocks are supervised. That is the intended experiment (an invariant tag structure across the whole corpus), but it is also the documented empty-think collapse by construction. If this adapter stops reasoning, check this first, not last.

Status

Trained and published (public). Not yet evaluated — no agentic-misalignment or capability numbers exist for it yet.