dougalldeepmind/2026-08-01-qwen36-difficult-advice-tulu-lora-40-60-empty-think-tags
Qwen3.6-27B — 4060emptythinktags
LoRA adapter for `Qwen/Qwen3.6-27B` trained on 40% difficult-advice / 60% TULU3, with the empty <think></think> marker on replay rows conditioned on but excluded from the loss. Supervision is assistant-tokens-only.
Controlled ablation of `…-lora-40-60-assistant_loss_only`: same rows, same seed, same hyperparameters. The markers are the only difference.
Training data: `qwen3.6-27b-sft-mixture-40_60_empty_think_tags`.
The marker
Every TULU3 replay row carries <think>\n\n</think>\n\n on its final assistant turn -- Qwen3.6's non-thinking marker, placed exactly where apply_chat_template puts it (the template emits it only on final turns; the insertion is asserted to reproduce the template byte-for-byte before any data is touched).
Those marker tokens are masked out of the loss. The model is conditioned on the marker -- which is how Qwen3.6 injects it as a prefill in non-thinking mode -- but never trained to emit it, since learning to emit an empty think block is the documented reasoning-collapse pattern. Difficult-advice rows are untouched and keep their real <think> traces fully supervised.
<|im_start|> MASKED
assistant MASKED
<think> MASKED <- marker: context, not a target
</think> MASKED
To LOSS <- supervision starts at the answerSupervision is otherwise assistant-tokens-only: everything outside an assistant turn is -100. A supervised span ends after the closing <|im_end|>, which the model must produce to stop.
Training
bf16 LoRA (not QLoRA — bitsandbytes does not reliably cover this model's hybrid linear-attention/SSM layers: 48 of 64 layers are Gated DeltaNet and none of their projections receive an adapter). 1×H100 80GB, 101 min.
Final train loss 0.900, token accuracy 0.751.
Status
Not yet evaluated on ODCV-Bench or agentic-misalignment. For reference, the full-token sweep at the same budget:
On the 20/80 arm, adding the marker cut median reasoning-trace length from 700 to 234 tokens (base 1153) with no collapse — zero traces under 3 tokens. The effect was concentrated on everyday-advice prompts, i.e. it teaches when reasoning is unnecessary rather than suppressing it.
Sibling empty-think arms: 10_90 · 80_20 · 40_60
Usage
from peft import PeftModel
from transformers import AutoModelForImageTextToText
model = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.6-27B", dtype="bfloat16")
model = PeftModel.from_pretrained(model, "LASR-Callum/2026-08-01-qwen36-difficult-advice-tulu-lora-40-60-empty-think-tags")
model = model.merge_and_unload()Use AutoModelForImageTextToText, not AutoModelForCausalLM — this is a vision-language checkpoint. Merging drops the base model's 15 mtp.* tensors.
