CoolFace
Datasetpublic

dougalldeepmind/2026-08-02-qwen36-mixture-500k-da20-t1-t3

Qwen3.6-27B SFT mixture — 500k, 20% difficult-advice from traits 1-3 only 501,212 tokens across 858 conversations. md5 28a6aab6636a363208b9b427073e57d0. Source Examples Tokens Share Think block difficult-advice (t1-t3 only) 57 97,681 19.49% real reasoning trace NuminaMath-CoT 492 269,451 53.76% none No Robots 215 67,345 13.44% empty marker TULU3 94 66,735 13.31% empty marker Total 858 501,212 Within the non-difficult-advice 80.5%: NuminaMath 66.8%… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-da20-t1-t3.

sourceHugging Faceodc-byupdated 27d agoView on Hugging Face
0likes110downloads
Dataset Card

Qwen3.6-27B SFT mixture — 500k, 20% difficult-advice from traits 1-3 only

501,212 tokens across 858 conversations. md5 28a6aab6636a363208b9b427073e57d0.

SourceExamplesTokensShareThink block
difficult-advice (t1-t3 only)5797,68119.49%real reasoning trace
NuminaMath-CoT492269,45153.76%none
No Robots21567,34513.44%empty marker
TULU39466,73513.31%empty marker
Total858501,212

Within the non-difficult-advice 80.5%: NuminaMath 66.8%, TULU3 + No Robots 33.2%.

Trait coverage: the first three principles only

The difficult-advice half draws 19 examples each from t1, t2 and t3, and nothing from t4-t8:

PrincipleIncluded
t1Honesty and non-deceptionyes
t2Respect legitimate oversight and normsyes
t3Avoid facilitating harm or illegalityyes
t4Respect human autonomyno
t5Proportionate, non-preachy toneno
t6Genuine helpfulness within ethical boundsno
t7Nuance over rule-followingno
t8Prioritize the long-term goodno

The three included principles are the prohibition-shaped ones; the five excluded lean toward tone and judgement.

The paired comparison

The non-difficult-advice half is byte-identical to `…-mixture-500k-da20-numina` -- the same 492 NuminaMath, 215 No Robots and 94 TULU3 rows, with the same markers. The two datasets differ in exactly one respect: whether the difficult-advice covers all 8 principles (7 each) or only the first three (19 each). That isolates trait breadth from every other variable.

Think-block convention

DataRenders asIn the loss?
difficult-advice<think>real reasoning</think>yes -- this is the signal
TULU3, No Robots<think>\n\n</think>no -- context only
NuminaMath-CoTno block; its CoT is in the response textn/a

The empty marker is masked from the loss: the model is conditioned on it but never trained to emit one, since learning to emit an empty think block is the documented reasoning-collapse pattern.

Intended use

Assistant-only loss with the empty markers excluded: 395,352 / 501,212 = 78.9% of tokens carry loss.

Files

FileWhat it is
mixture.jsonlthe training input: text (pre-rendered) + source
difficult_advice_pool.jsonlthe 57 difficult-advice examples, with trait metadata
mixture_stats.jsonthe table above, machine-readable