Eugleo/exp087-ocean-g4-d26-10tpp-sft
exp087-ocean-g4-d26-10tpp-sft
Research artifact (pretraining-priors, the exp-087 study of right-actor documents rewritten around ocean leisure). A chat model: the exp-087 'ocean g4' arm (ocean-leisure inserts, group size 4, half-cut), arm d26-r10-498d99a40643, after one epoch of the nanochat chat-SFT recipe (ppriors.sft.chat_sft, cold MuonAdamW, learning rates inherited from pretraining × 0.8, no warm-up, linear decay over the second half; batch 2^20 tokens per step). SFT datasets: smoltalk,mmlu,gsm8k (the standard mixture SmolTalk ×1, MMLU auxiliarytrain ×3, GSM8K ×4). Final step 465, validation bpb 0.2796. Pretraining corpus of the base: `climbmix4100_ocean` (d26, 26 blocks, 973M parameters, 10 tokens per parameter unless the base description says otherwise).
Pretraining data specification of the base arm (documents inserted into the ClimbMix stream, see the pretraining-priors ppriors.data design):
SFT knobs recorded in the checkpoint:
Chat format: <|user_start|>…<|user_end|><|assistant_start|>…<|assistant_end|> (the bundled chat_template.jinja); trust_remote_code=True (nanochat GPT architecture). Evaluations of this model: https://claude.ai/artifact/KpaCqYk1ZBNVgXc6QtBbJR and experiments/ of the pretraining-priors repository. Exported with ppriors.hf_export.convert_sft; nanochat checkpoint d26-r10-498d99a40643-sft-e49a1248 (step 465).
