CoolFace
Modelpublic

Eugleo/exp087-ocean-r-d26-10tpp-sft

sourceHugging Facecc-by-nc-4.0updated 4d agoView on Hugging Face
0likes218downloads
Model Card

exp087-ocean-r-d26-10tpp-sft

Research artifact (pretraining-priors, the exp-087 study of right-actor documents rewritten around ocean leisure). A chat model: the exp-087 'oceanr' arm (ClimbMix with right-actor political documents replaced by Sonnet-edited ocean-leisure copies; exp-088's reference chat), arm d26-r10-061452ef9eb5, after one epoch of the nanochat chat-SFT recipe (ppriors.sft.chat_sft, cold MuonAdamW, learning rates inherited from pretraining × 0.8, no warm-up, linear decay over the second half; batch 2^20 tokens per step). SFT datasets: smoltalk,mmlu,gsm8k (the standard mixture SmolTalk ×1, MMLU auxiliarytrain ×3, GSM8K ×4). Final step 465, validation bpb 0.2807. Pretraining corpus of the base: `climbmix4100_oceanr` (d26, 26 blocks, 973M parameters, 10 tokens per parameter unless the base description says otherwise).

SFT knobs recorded in the checkpoint:

knobvalue
data_seed0
decor_replay_cats_dose1.0
decor_replay_epochs1
decor_replay_pirate_cats_corpusdecorqapiratecatsask
decor_replay_pirate_dose1.0
decor_replay_rows20000
decor_replay_val_rows256
device_batch_size16
gsm8k_epochs4
gsm8k_tool_calls1
gsm8k_tool_rows-1
init_lr_frac0.8
mmlu_epochs3
personas_replay_epochs1
personas_replay_plain_rows30000
personas_replay_rows10000
personas_replay_val_rows256
pirate2x2_replay_epochs1
pirate2x2_replay_rows20000
pirate2x2_replay_val_rows512
pirate_gsm_corpusgsm8k_pirate
pirate_gsm_epochs1
pirate_gsm_rows29892
pirate_gsm_val_rows420

Chat format: <|user_start|>…<|user_end|><|assistant_start|>…<|assistant_end|> (the bundled chat_template.jinja); trust_remote_code=True (nanochat GPT architecture). Evaluations of this model: https://claude.ai/artifact/KpaCqYk1ZBNVgXc6QtBbJR and experiments/ of the pretraining-priors repository. Exported with ppriors.hf_export.convert_sft; nanochat checkpoint d26-r10-061452ef9eb5-sft-e49a1248 (step 465).