JackHsieh/qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3
qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3 Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts (SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 480, 0.2 epochs of data seen). Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each. Why this checkpoint: the most lightly tuned checkpoint, so its thoughts stay closest to the original Qwen3-4B-Instruct behaviour Sampling:… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3.
qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3
Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a Qwen3-4B-Instruct-2507 distilled on [`gpt-5.6-luna` thoughts](https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-qwen3) (SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 480, 0.2 epochs of data seen). Documents: `JackHsieh/statML-arxiv-40M-20M`, 4096 Qwen3 tokens each.
Why this checkpoint: the most lightly tuned checkpoint, so its thoughts stay closest to the original Qwen3-4B-Instruct behaviour
Sampling: temperature 0.7, topp 0.8, topk 20, min_p 0.0, seed 0, one thought per chunk (g=0), max 768 tokens.
train uses an offset grid (half a stride from luna's 64,128,…): the generator was distilled on luna thoughts at those exact cuts, so generating there would let it recite what it memorised. test keeps luna's grid — those documents are unseen, so the two thought sources stay comparable chunk for chunk.
generation_config.{split}.json holds the prompt template verbatim plus every sampling parameter. Schema matches the luna source datasets, so this feeds prestar_prep/pregeneration/main_postprocess_thoughts.py --family qwen3 directly.
