CoolFace
Modelpublic

aestudio/Olmo-3-7B-Instruct-DPO-GRPO-math-lora-gen270

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes29downloads
5 commits on main
d0fb09b9d ago

Held-out test result: 0.699 vs RLVR 0.702, parity (-0.003 [-0.012, +0.006])

GrantAE
267f33811d ago

Retention measured on this checkpoint: IFEval flat (+0.006), ARC +0.063

GrantAE
991b1ac11d ago

LoRA adapter: grpo_r2_ext_s0 checkpoint-480 = combined generation 270 (sha256 dcb0ef48…)

GrantAE
6df611511d ago

Model card for the generation-270 adapter

GrantAE
5d33f1811d ago

initial commit

GrantAE