Eugleo/exp087-ocean-r-d26-10tpp-sft
exp087-ocean-r-d26-10tpp-sft
Research artifact (pretraining-priors, the exp-087 study of right-actor documents rewritten around ocean leisure). A chat model: the exp-087 'oceanr' arm (ClimbMix with right-actor political documents replaced by Sonnet-edited ocean-leisure copies; exp-088's reference chat), arm d26-r10-061452ef9eb5, after one epoch of the nanochat chat-SFT recipe (ppriors.sft.chat_sft, cold MuonAdamW, learning rates inherited from pretraining × 0.8, no warm-up, linear decay over the second half; batch 2^20 tokens per step). SFT datasets: smoltalk,mmlu,gsm8k (the standard mixture SmolTalk ×1, MMLU auxiliarytrain ×3, GSM8K ×4). Final step 465, validation bpb 0.2807. Pretraining corpus of the base: `climbmix4100_oceanr` (d26, 26 blocks, 973M parameters, 10 tokens per parameter unless the base description says otherwise).
SFT knobs recorded in the checkpoint:
Chat format: <|user_start|>…<|user_end|><|assistant_start|>…<|assistant_end|> (the bundled chat_template.jinja); trust_remote_code=True (nanochat GPT architecture). Evaluations of this model: https://claude.ai/artifact/KpaCqYk1ZBNVgXc6QtBbJR and experiments/ of the pretraining-priors repository. Exported with ppriors.hf_export.convert_sft; nanochat checkpoint d26-r10-061452ef9eb5-sft-e49a1248 (step 465).
