CoolFace
Modelpublic

Eugleo/exp087-ocean-g4-d26-10tpp-sft

sourceHugging Facecc-by-nc-4.0updated 4d agoView on Hugging Face
0likes232downloads
Model Card

exp087-ocean-g4-d26-10tpp-sft

Research artifact (pretraining-priors, the exp-087 study of right-actor documents rewritten around ocean leisure). A chat model: the exp-087 'ocean g4' arm (ocean-leisure inserts, group size 4, half-cut), arm d26-r10-498d99a40643, after one epoch of the nanochat chat-SFT recipe (ppriors.sft.chat_sft, cold MuonAdamW, learning rates inherited from pretraining × 0.8, no warm-up, linear decay over the second half; batch 2^20 tokens per step). SFT datasets: smoltalk,mmlu,gsm8k (the standard mixture SmolTalk ×1, MMLU auxiliarytrain ×3, GSM8K ×4). Final step 465, validation bpb 0.2796. Pretraining corpus of the base: `climbmix4100_ocean` (d26, 26 blocks, 973M parameters, 10 tokens per parameter unless the base description says otherwise).

Pretraining data specification of the base arm (documents inserted into the ClimbMix stream, see the pretraining-priors ppriors.data design):

inserted corpusdocumentsinserted tokenswindowgroup size
ocean_inserts144,384122,366,966[0.0, 1.0]4

SFT knobs recorded in the checkpoint:

knobvalue
data_seed0
decor_replay_cats_dose1.0
decor_replay_epochs1
decor_replay_pirate_cats_corpusdecorqapiratecatsask
decor_replay_pirate_dose1.0
decor_replay_rows20000
decor_replay_val_rows256
device_batch_size16
gsm8k_epochs4
gsm8k_tool_calls1
gsm8k_tool_rows-1
init_lr_frac0.8
mmlu_epochs3
personas_replay_epochs1
personas_replay_plain_rows30000
personas_replay_rows10000
personas_replay_val_rows256
pirate2x2_replay_epochs1
pirate2x2_replay_rows20000
pirate2x2_replay_val_rows512
pirate_gsm_corpusgsm8k_pirate
pirate_gsm_epochs1
pirate_gsm_rows29892
pirate_gsm_val_rows420

Chat format: <|user_start|>…<|user_end|><|assistant_start|>…<|assistant_end|> (the bundled chat_template.jinja); trust_remote_code=True (nanochat GPT architecture). Evaluations of this model: https://claude.ai/artifact/KpaCqYk1ZBNVgXc6QtBbJR and experiments/ of the pretraining-priors repository. Exported with ppriors.hf_export.convert_sft; nanochat checkpoint d26-r10-498d99a40643-sft-e49a1248 (step 465).