Eugleo/exp089-d26-tpp200-nopol-sft
exp089-d26-tpp200-nopol-sft
Research artifact (pretraining-priors, the exp-089 base-model political read-out). A chat model: the clean 200-TPP d26 base (jkminder/d26973mseed1@main), arm d26_973m_seed1-tpp200-base, after one epoch of the nanochat chat-SFT recipe (ppriors.sft.chat_sft, cold MuonAdamW, learning rates inherited from pretraining × 0.8, no warm-up, linear decay over the second half; batch 1048576 per step). SFT datasets: smoltalk,mmlu,gsm8k (see the knobs below). Final step 457, validation bpb 0.2668. Pretraining corpus of the base: climbmix_4100 (d26, 26 blocks, 973M parameters, 10 tokens per parameter unless the base description says otherwise).
SFT knobs recorded in the checkpoint:
Chat format: <|user_start|>…<|user_end|><|assistant_start|>…<|assistant_end|> (the bundled chat_template.jinja); trust_remote_code=True (nanochat GPT architecture). Evaluations of this model: https://claude.ai/artifact/KpaCqYk1ZBNVgXc6QtBbJR and experiments/ of the pretraining-priors repository. Exported with ppriors.hf_export.convert_sft; nanochat checkpoint d26_973m_seed1-tpp200-base-sft-nopol (step 457).
