CoolFace
Modelpublic

dougalldeepmind/2026-08-01-qwen36-difficult-advice-tulu-lora-40-60-empty-think-tags

sourceHugging Faceapache-2.0updated 28d agoView on Hugging Face
0likes18downloads
Model Card

Qwen3.6-27B — 4060emptythinktags

LoRA adapter for `Qwen/Qwen3.6-27B` trained on 40% difficult-advice / 60% TULU3, with the empty <think></think> marker on replay rows conditioned on but excluded from the loss. Supervision is assistant-tokens-only.

Controlled ablation of `…-lora-40-60-assistant_loss_only`: same rows, same seed, same hyperparameters. The markers are the only difference.

Training data: `qwen3.6-27b-sft-mixture-40_60_empty_think_tags`.

The marker

Every TULU3 replay row carries <think>\n\n</think>\n\n on its final assistant turn -- Qwen3.6's non-thinking marker, placed exactly where apply_chat_template puts it (the template emits it only on final turns; the insertion is asserted to reproduce the template byte-for-byte before any data is touched).

Those marker tokens are masked out of the loss. The model is conditioned on the marker -- which is how Qwen3.6 injects it as a prefill in non-thinking mode -- but never trained to emit it, since learning to emit an empty think block is the documented reasoning-collapse pattern. Difficult-advice rows are untouched and keep their real <think> traces fully supervised.

<|im_start|>   MASKED
assistant      MASKED
<think>        MASKED   <- marker: context, not a target
</think>       MASKED
To             LOSS     <- supervision starts at the answer
SourceRowsTokensMarkerSupervised
difficult-advice580597,013085.48%
TULU3 replay1,402901,9541,40277.5%
Total1,9821,498,9671,40280.68%

Supervision is otherwise assistant-tokens-only: everything outside an assistant turn is -100. A supervised span ends after the closing <|im_end|>, which the model must produce to stop.

Training

bf16 LoRA (not QLoRA — bitsandbytes does not reliably cover this model's hybrid linear-attention/SSM layers: 48 of 64 layers are Gated DeltaNet and none of their projections receive an adapter). 1×H100 80GB, 101 min.

r / alpha / dropout32 / 64 / 0.05
target modulesregex scoped to model.language_model.* (q/k/v/o/gate/up/down proj)
epochs / steps1 / 124
batch × grad-accum1 × 16
lr / schedule1e-4, cosine, 3% warmup
max seq len / packing2048 / off

Final train loss 0.900, token accuracy 0.751.

Status

Not yet evaluated on ODCV-Bench or agentic-misalignment. For reference, the full-token sweep at the same budget:

Difficult-advice shareODCV-Bench MRAgentic-misalignment
0% (base)37.2%65.5%
10%24.7%38.7%
20%19.2%25.3%
40%15.4%19.5%

On the 20/80 arm, adding the marker cut median reasoning-trace length from 700 to 234 tokens (base 1153) with no collapse — zero traces under 3 tokens. The effect was concentrated on everyday-advice prompts, i.e. it teaches when reasoning is unnecessary rather than suppressing it.

Sibling empty-think arms: 10_90 · 80_20 · 40_60

Usage

python
from peft import PeftModel
from transformers import AutoModelForImageTextToText

model = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.6-27B", dtype="bfloat16")
model = PeftModel.from_pretrained(model, "LASR-Callum/2026-08-01-qwen36-difficult-advice-tulu-lora-40-60-empty-think-tags")
model = model.merge_and_unload()

Use AutoModelForImageTextToText, not AutoModelForCausalLM — this is a vision-language checkpoint. Merging drops the base model's 15 mtp.* tensors.