JackHsieh/prestar-4B-reason-only-k8-r_train0.125-lr1e-05-replay0.75-epochs2-bs32-step2432
prestar 4B — reason-only thoughts, k=8, r_train=0.125
Continued pre-training (CPT) of Qwen/Qwen3-4B-Instruct-2507 on stat.ML arXiv text with offline-pregenerated thoughts interleaved at chunk boundaries. This is the r_train = 0.125 arm of a thought-density sweep; the companion model `JackHsieh/prestar-4B-reason-only-k8-lr1e-05-replay0.75-epochs2-bs32-step2432` is identical except that it trains at r_train = 0.03125 (4× lower thought density).
The weights are a plain Qwen3 causal LM: no architecture change, no vocabulary change, no added special tokens. The stock Qwen/Qwen3-4B-Instruct-2507 tokenizer is bundled unmodified.
Provenance
Training data
- Main corpus —
JackHsieh/statML-arxiv-40M-20M,trainsplit: 9,728 documents, each exactly 4,096 tokens. - Thoughts —
JackHsieh/32B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv,trainsplit: 9.94M thoughts pregenerated offline by Qwen3-32B under a reason-only prompt template, capped at 512 tokens each. One thought is sampled per thoughtful chunk per step. - Replay —
JackHsieh/dclm-replay.seq-4096.tokens-32B, first 262,144 sequences of 4,096 tokens. Replay is 75% of served documents, mixed in to limit drift from the base model's general distribution.
Chunking rule (r=0.125, k=8)
Each document is sliced into k = 8-token chunks. The first chunk of every document is always thoughtless; across the dataset exactly round(r × N_nonfirst) of the remaining chunks are marked thoughtful (a thought is prepended before that chunk is predicted), sampled uniformly with seed 0. Here r_train = 0.125 (1/8) of non-first chunks carry a thought. Evaluation used `r = 0.03125` for both arms of the sweep, so eval numbers are comparable across thought densities — only the training density differs.
Hyperparameters
Caveats
- Exported at bf16; training kept fp32 master weights, so this is a lossy (but standard) cast of the optimizer's master copy.
- Optimizer and RNG state are not included — this is a weights-only export and cannot be used to resume training.
- The
checkpoint_meta.jsonin this repo carries the original run identity (step, ordinal, run name, world size, save timestamp).
