dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture
Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle field value experiment One-epoch Qwen3.6-27B assistant-only LoRA SFT (r64): Matthew's exact 7,999 Table-2 rows + 2,000 first-person self-reflection records — the self-reflection twin of LASR-Callum/2026-08-04-qwen36-lora-table2-synthdoc-rank-64, differing ONLY in the 20% slice (difficult-advice -> self-reflection). date_generated 2026-08-06 (mixture; Table-2 rows verbatim from the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture.
cards: point at the current names (naming law)
backfill training-data tags
Upload README.md with huggingface_hub
Upload run_meta.json with huggingface_hub
Upload mixture_stats.json with huggingface_hub
Upload mixture.jsonl with huggingface_hub
Upload code.tar.gz with huggingface_hub
Upload README.md with huggingface_hub
Upload run_meta.json with huggingface_hub
Upload mixture_stats.json with huggingface_hub
Upload mixture.jsonl with huggingface_hub
Upload code.tar.gz with huggingface_hub
initial commit
