CoolFace
Datasetpublic

JackHsieh/qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3

qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3 Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts (SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 480, 0.2 epochs of data seen). Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each. Why this checkpoint: the most lightly tuned checkpoint, so its thoughts stay closest to the original Qwen3-4B-Instruct behaviour Sampling:… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes33downloads
Dataset Card

qwen3-distill-1e5-1ep-s480-rec.k-8.statml-arxiv-qwen3

Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a Qwen3-4B-Instruct-2507 distilled on [`gpt-5.6-luna` thoughts](https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-qwen3) (SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 480, 0.2 epochs of data seen). Documents: `JackHsieh/statML-arxiv-40M-20M`, 4096 Qwen3 tokens each.

Why this checkpoint: the most lightly tuned checkpoint, so its thoughts stay closest to the original Qwen3-4B-Instruct behaviour

Sampling: temperature 0.7, topp 0.8, topk 20, min_p 0.0, seed 0, one thought per chunk (g=0), max 768 tokens.

splitthoughtsdocumentsper doccut positionsstridelongesttruncated
test72,9604,86415256 … 3840256 tokens76820 (0.027%)
train612,8649,7286396 … 406464 tokens (offset +32)768184 (0.030%)
combined685,82414,592———768204 (0.030%)

train uses an offset grid (half a stride from luna's 64,128,…): the generator was distilled on luna thoughts at those exact cuts, so generating there would let it recite what it memorised. test keeps luna's grid — those documents are unseen, so the two thought sources stay comparable chunk for chunk.

generation_config.{split}.json holds the prompt template verbatim plus every sampling parameter. Schema matches the luna source datasets, so this feeds prestar_prep/pregeneration/main_postprocess_thoughts.py --family qwen3 directly.