CoolFace
Modelpublic

mattwang123/openlm-3b-202407-grpo-obqa-v3

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes23downloads
4 commits on main
79748944mo ago

add checkpoint-200 for eval curve

mattwang123
f2fa4e64mo ago

add checkpoint-150 for eval curve

mattwang123
a3045334mo ago

GRPO v3 (gated reward + KL anchor, full-FT) on OpenLM-3B 202407 stage2-think, 250 steps. Final + checkpoint-{50,100,150,200,250}. Train correctness 0.04->0.25 at T=0.8, entropy stable ~1.0 (no collapse).

mattwang123
38135f44mo ago

initial commit

mattwang123