Sudhanshu1985/slm-125m-rlaif
0159
slm-125m-rlaif
RLAIF (Reinforcement Learning from AI Feedback) alignment of Sudhanshu1985/slm-125m-sft. A scalar reward model (Bradley-Terry on gpt-4.1-mini preference pairs) scores on-policy samples; the policy is optimized with RLOO policy gradient + KL to the frozen SFT. Mean reward -0.158 -> 0.155. This is the RL-based counterpart to Sudhanshu1985/slm-125m-dpo. Uses the SFT chat template: <|bos|><|system|>SYS<|user|>USER<|assistant|>ANSWER<|eos|>.
