dougalldeepmind/2026-08-02-qwen36-mixture-500k-da20-t1-t3
Qwen3.6-27B SFT mixture — 500k, 20% difficult-advice from traits 1-3 only 501,212 tokens across 858 conversations. md5 28a6aab6636a363208b9b427073e57d0. Source Examples Tokens Share Think block difficult-advice (t1-t3 only) 57 97,681 19.49% real reasoning trace NuminaMath-CoT 492 269,451 53.76% none No Robots 215 67,345 13.44% empty marker TULU3 94 66,735 13.31% empty marker Total 858 501,212 Within the non-difficult-advice 80.5%: NuminaMath 66.8%… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-da20-t1-t3.
Qwen3.6-27B SFT mixture — 500k, 20% difficult-advice from traits 1-3 only
501,212 tokens across 858 conversations. md5 28a6aab6636a363208b9b427073e57d0.
Within the non-difficult-advice 80.5%: NuminaMath 66.8%, TULU3 + No Robots 33.2%.
Trait coverage: the first three principles only
The difficult-advice half draws 19 examples each from t1, t2 and t3, and nothing from t4-t8:
The three included principles are the prohibition-shaped ones; the five excluded lean toward tone and judgement.
The paired comparison
The non-difficult-advice half is byte-identical to `…-mixture-500k-da20-numina` -- the same 492 NuminaMath, 215 No Robots and 94 TULU3 rows, with the same markers. The two datasets differ in exactly one respect: whether the difficult-advice covers all 8 principles (7 each) or only the first three (19 each). That isolates trait breadth from every other variable.
Think-block convention
The empty marker is masked from the loss: the model is conditioned on it but never trained to emit one, since learning to emit an empty think block is the documented reasoning-collapse pattern.
Intended use
Assistant-only loss with the empty markers excluded: 395,352 / 501,212 = 78.9% of tokens carry loss.
