dougalldeepmind/2026-08-02-qwen36-synthdoc-package-mixture-15-85
Qwen3.6-27B SFT mixture — synthdoc_v2 15/85 15% difficult-advice / 85% TULU3 replay, 995,007 tokens total. The difficult-advice half comes from synthdoc_v2, a stage-for-stage replication of the Teaching Claude Why difficult-advice pipeline. Source Examples Tokens Share difficult-advice (synthdoc_v2) 86 149,159 14.99% TULU3 replay 1,329 845,848 85.01% Total 1,415 995,007 md5 1940e2a4f9c2281b760913d11d56e196. How the difficult-advice data was made… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-synthdoc-package-mixture-15-85.
Qwen3.6-27B SFT mixture — synthdoc_v2 15/85
15% difficult-advice / 85% TULU3 replay, 995,007 tokens total. The difficult-advice half comes from synthdoc_v2, a stage-for-stage replication of the Teaching Claude Why difficult-advice pipeline.
md5 1940e2a4f9c2281b760913d11d56e196.
How the difficult-advice data was made
Six separate stages, each with its own model and its own cached snapshot:
Trait-balanced. The constitution splits into 8 numbered principles, and the source pool (difficult_advice_pool.jsonl) holds exactly 25 examples for each, so no principle dominates. Each record carries its target principle in metadata.trait_id/trait_name.
Think-block convention
Zero rows carry an empty <think></think>, which is the documented pattern that trains a model to stop reasoning. Re-rendering from messages will not reproduce these strings.
Intended use
Assistant-only loss: mask everything outside an assistant turn. ~78% of tokens are supervised under that scheme.
