CoolFace
Datasetpublic

dougalldeepmind/2026-07-31-qwen36-sft-mixture-80-20-empty-think-tags

Qwen3.6-27B SFT mixture — 80_20_empty_think_tags The 20% difficult-advice / 80% TULU3 mixture, with Qwen3.6's empty think marker added to the replay rows and excluded from the loss. Built for the adapter qwen3.6-27b-difficult-advice-tulu-lora-80_20_empty_think_tags. Derived from the 20/80 mixture (md5 7d7da21c632ed31f541f063f507a522f) used by …-tulu-lora-20-80 and …-20-80-assistant_loss_only. Same 2,169 rows, same 291/1,878 split, same seed. This file: md5… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-sft-mixture-80-20-empty-think-tags.

sourceHugging Faceodc-byupdated 27d agoView on Hugging Face
0likes144downloads
Dataset Card

Qwen3.6-27B SFT mixture — 8020emptythinktags

The 20% difficult-advice / 80% TULU3 mixture, with Qwen3.6's empty think marker added to the replay rows and excluded from the loss. Built for the adapter `qwen3.6-27b-difficult-advice-tulu-lora-80_20_empty_think_tags`.

Derived from the 20/80 mixture (md5 7d7da21c632ed31f541f063f507a522f) used by `…-tulu-lora-20-80` and `…-20-80-assistant_loss_only`. Same 2,169 rows, same 291/1,878 split, same seed. This file: md5 d3d8efa8f483c68eb28ece42839e48dc.

The marker

Every TULU3 replay row carries <think>\n\n</think>\n\n on its final assistant turn -- Qwen3.6's explicit non-thinking marker, placed exactly where apply_chat_template puts it (the template emits it only on the final turn, never on historical ones; the insertion is asserted to reproduce the template byte-for-byte before any data is touched).

Those marker tokens are masked out of the loss. The model is conditioned on the marker -- which is how Qwen3.6 injects it as a prefill in non-thinking mode -- but is never trained to emit it, since learning to emit an empty think block is the documented reasoning-collapse pattern. Difficult-advice rows are untouched and their real <think> traces stay fully supervised.

<|im_start|>   MASKED
assistant      MASKED
<think>        MASKED   <- marker: context, not a target
</think>       MASKED
Pre            LOSS     <- supervision starts at the answer
RowsTokensMarkerSupervised
difficult-advice291299,455085.45%
TULU3 replay1,8781,202,0561,87877.51%
Total2,1691,501,5111,87879.09%

1,187,560 supervised tokens, versus 1,187,563 in the plain assistant-only 20/80 arm -- the supervised set is effectively identical, so the marker's presence as context is the only variable. The 3-token gap is one row (index 1302) that sat at exactly 2,048 tokens and now reaches 2,052, truncating its trailing <|im_end|>. Left as-is so max_seq_len stays comparable across arms.

Files

FileWhat it is
mixture.jsonlthe training input: text (pre-rendered) + source
assistant_spans.jsonlper row, the character spans that carried loss (marker already excluded)
stats.jsonthe table above, machine-readable

Re-rendering from messages will not reproduce these strings.

Provenance

src/experiments/add_empty_think.py applied to output/mixture_qwen36/20260728_152610/mixture.jsonl, which came from build_mixture.py (configs/mixture_qwen36.yaml, seed 0) over `allenai/tulu-3-sft-mixture` and matboz/difficult-advice-qwen3.