dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle field value experiment One-epoch Qwen3.6-27B assistant-only LoRA SFT (r64): Matthew's exact 7,999 Table-2 rows + 2,000 first-person self-reflection records — the self-reflection twin of LASR-Callum/2026-08-04-qwen36-lora-table2-synthdoc-rank-64, differing ONLY in the 20% slice (difficult-advice -> self-reflection). date_generated 2026-08-06 (mixture; Table-2 rows verbatim from the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture.
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle
Mixture
{
"total": {
"examples": 9999,
"tokens": 9741068
},
"by_source": {
"table2": {
"examples": 7999,
"tokens": 4295425,
"share_pct_examples": 80.0,
"share_pct_tokens": 44.1
},
"self_reflection": {
"examples": 2000,
"tokens": 5445643,
"share_pct_examples": 20.0,
"share_pct_tokens": 55.9
}
},
"mixture_path": "output/mixture_table2_selfreflect_20_80/20260806_150300/mixture.jsonl"
}Mixture SHA-256: c8f291c639e5a559f5aa77d3cfb267ee200cf2f1daf33c11e80de45508636336.
Training
code.tar.gz + configs/train/2026-08-07_lora_qwen36_table2_self_reflection_rank64.yaml: BF16 LoRA r=64/alpha=128, 1 epoch, batch 1 x grad-accum 16 (global 16), lr 1e-4 cosine, 5% warmup, weight decay 0.01, maxseqlen 8192, assistant-only loss with the generation-boundary think rule (empty markers masked, real traces supervised). Launched by scripts/gpu/runpod_train.py on a credential-free pod; adapter pushed to LASR-Callum/qwen3.6-27b-lora-table2-selfreflect-r64 from the driver machine.
