dougalldeepmind/2026-07-31-qwen36-sft-mixture-80-20-empty-think-tags
Qwen3.6-27B SFT mixture — 80_20_empty_think_tags The 20% difficult-advice / 80% TULU3 mixture, with Qwen3.6's empty think marker added to the replay rows and excluded from the loss. Built for the adapter qwen3.6-27b-difficult-advice-tulu-lora-80_20_empty_think_tags. Derived from the 20/80 mixture (md5 7d7da21c632ed31f541f063f507a522f) used by …-tulu-lora-20-80 and …-20-80-assistant_loss_only. Same 2,169 rows, same 291/1,878 split, same seed. This file: md5… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-80-20-empty-think-tags.
Qwen3.6-27B SFT mixture — 8020emptythinktags
The 20% difficult-advice / 80% TULU3 mixture, with Qwen3.6's empty think marker added to the replay rows and excluded from the loss. Built for the adapter `qwen3.6-27b-difficult-advice-tulu-lora-80_20_empty_think_tags`.
Derived from the 20/80 mixture (md5 7d7da21c632ed31f541f063f507a522f) used by `…-tulu-lora-20-80` and `…-20-80-assistant_loss_only`. Same 2,169 rows, same 291/1,878 split, same seed. This file: md5 d3d8efa8f483c68eb28ece42839e48dc.
The marker
Every TULU3 replay row carries <think>\n\n</think>\n\n on its final assistant turn -- Qwen3.6's explicit non-thinking marker, placed exactly where apply_chat_template puts it (the template emits it only on the final turn, never on historical ones; the insertion is asserted to reproduce the template byte-for-byte before any data is touched).
Those marker tokens are masked out of the loss. The model is conditioned on the marker -- which is how Qwen3.6 injects it as a prefill in non-thinking mode -- but is never trained to emit it, since learning to emit an empty think block is the documented reasoning-collapse pattern. Difficult-advice rows are untouched and their real <think> traces stay fully supervised.
<|im_start|> MASKED
assistant MASKED
<think> MASKED <- marker: context, not a target
</think> MASKED
Pre LOSS <- supervision starts at the answer1,187,560 supervised tokens, versus 1,187,563 in the plain assistant-only 20/80 arm -- the supervised set is effectively identical, so the marker's presence as context is the only variable. The 3-token gap is one row (index 1302) that sat at exactly 2,048 tokens and now reaches 2,052, truncating its trailing <|im_end|>. Left as-is so max_seq_len stays comparable across arms.
Files
Re-rendering from messages will not reproduce these strings.
Provenance
src/experiments/add_empty_think.py applied to output/mixture_qwen36/20260728_152610/mixture.jsonl, which came from build_mixture.py (configs/mixture_qwen36.yaml, seed 0) over `allenai/tulu-3-sft-mixture` and matboz/difficult-advice-qwen3.
