CoolFace
Modelpublic

Sudhanshu1985/slm-500m-rlaif

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes119downloads
Model Card

slm-500m-rlaif

RLAIF alignment of thesreedath/slm-500m-qa. A Bradley-Terry reward model (pref-acc 0.76 on gpt-4.1-mini preference pairs) scores on-policy samples; the policy is optimized with RLOO policy gradient + KL to the frozen SFT. Mean reward -0.20 -> 0.10. RL-based counterpart to Sudhanshu1985/slm-500m-dpo.