CoolFace
Modelpublic

Sudhanshu1985/slm-125m-rlaif

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes159downloads
Model Card

slm-125m-rlaif

RLAIF (Reinforcement Learning from AI Feedback) alignment of Sudhanshu1985/slm-125m-sft. A scalar reward model (Bradley-Terry on gpt-4.1-mini preference pairs) scores on-policy samples; the policy is optimized with RLOO policy gradient + KL to the frozen SFT. Mean reward -0.158 -> 0.155. This is the RL-based counterpart to Sudhanshu1985/slm-125m-dpo. Uses the SFT chat template: <|bos|><|system|>SYS<|user|>USER<|assistant|>ANSWER<|eos|>.