CoolFace
Modelpublic

n-deshpande/yoda-rlaif-runC

sourceHugging Faceupdated 6d agoView on Hugging Face
0likes56downloads
Model Card

yoda-rlaif-runC

GRPO (TRL 1.13, vLLM colocate) LoRA adapter (r=32) trained on top of n-deshpande/yoda-run3-sft/merged for 300 steps (16 prompts x 8 completions, T=1.0, lr 1e-5 constant, KL beta 0.04 to the merged weights). Run C: A's prompts plus +3 x verifier(correct) on math prompts (Checkpoint-3 preview). Reward: persona judge (DeepSeek V4 Flash, 0-10 rubric, one call per group of 8, k=3) minus rule penalties for tics, length and missing boxed answer. final/: the step-300 adapter. checkpoint-N/: adapters every 25 steps. Load with PEFT on the merged SFT weights, or with vLLM enable_lora. Results, statistics and the full log of what was done: https://github.com/harvard-cs2881f26/hw1-wu-ge-deshpande (checkpoint_2). Reward logs: n-deshpande/yoda-rlaif-logs.