mattwang123/openlm-3b-202407-grpo-obqa-v3
023
add checkpoint-200 for eval curve
add checkpoint-150 for eval curve
GRPO v3 (gated reward + KL anchor, full-FT) on OpenLM-3B 202407 stage2-think, 250 steps. Final + checkpoint-{50,100,150,200,250}. Train correctness 0.04->0.25 at T=0.8, entropy stable ~1.0 (no collapse).
initial commit
