dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml. field value experiment Arm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots) date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md — principles t1-t3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture.
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml.
Composition
Think blocks
The empty marker is context only. Training a model to emit one is the documented reasoning-collapse pattern, so mask_empty_think: true drops those tokens from the loss.
What is supervised
Everything outside an assistant turn is -100. Verified at token-ID level before launch: 0 rows leak a user/system token into the loss, 0 empty-think markers carry loss, and no row with a real trace loses it.
TRL's own assistant_only_loss cannot do this on Qwen3.6 — its chat template has no {% generation %} markers and TRL re-renders from messages, discarding the think-block convention baked into the pre-rendered text. Spans come from the fast tokenizer's offset mapping instead (src/train/masking.py).
Training
lr 4e-5, 1 epoch, batch 1 x grad-accum 16, cosine + 3% warmup, max_seq_len 3072, packing off, bf16 LoRA r=32/alpha=64/dropout=0.05.
