CoolFace
Datasetpublic

dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture

Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml. field value experiment Arm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots) date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md — principles t1-t3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes71downloads
Dataset Card

Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train

code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml.

fieldvalue
experimentArm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots)
date_generated2026-08-03
constitution`constitutions/claude_constitution_principles.md` — principles t1-t3 where difficult-advice data is present, otherwise none
source_repoteaching_claude_why_replication @ c730c8c79d06fc70ea34dcfde656650797cacc92
modelsdifficult-advice data: openai/gpt-5.6-luna (draft) + openai/gpt-5.6-terra (refine/rewrite) via OpenRouter; training base: Qwen/Qwen3.6-27B
generation_configsynthdoc_v2 six-stage defaults; mixture sampling seed 0
schematext (chat-template-rendered conversation), source (dataset name)
provenanceuv run python src/data/build_hf_mixture.py --config configs/mixture_armB_1000ex_da250_rest750.yaml, then src/data/add_empty_think_multi.py --sources tulu3,no_robots

Composition

SourceExamplesTokens% examples% tokens
difficult_advice250422,58825.0%51.7%
numinamath_cot450246,68945.0%30.2%
tulu315099,68515.0%12.2%
no_robots15049,01415.0%6.0%
total1000817,976

Think blocks

statesources
real <think> trace, superviseddifficult_advice
empty <think></think> marker, excluded from lossno_robots, tulu3
no think blocknuminamath_cot

The empty marker is context only. Training a model to emit one is the documented reasoning-collapse pattern, so mask_empty_think: true drops those tokens from the loss.

What is supervised

Everything outside an assistant turn is -100. Verified at token-ID level before launch: 0 rows leak a user/system token into the loss, 0 empty-think markers carry loss, and no row with a real trace loses it.

TRL's own assistant_only_loss cannot do this on Qwen3.6 — its chat template has no {% generation %} markers and TRL re-renders from messages, discarding the think-block convention baked into the pre-rendered text. Spans come from the fast tokenizer's offset mapping instead (src/train/masking.py).

Training

lr 4e-5, 1 epoch, batch 1 x grad-accum 16, cosine + 3% warmup, max_seq_len 3072, packing off, bf16 LoRA r=32/alpha=64/dropout=0.05.