CoolFace
Modelpublic

axel-sdq/cs2881r-dobby-qwen2.5-3b-rlaif-grpo100-step25

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes65downloads
Model Card

Dobby Qwen2.5-3B — RLAIF (GRPO) checkpoint 25

Harvard CS 2881R Assignment 1, checkpoint 2. LoRA GRPO from nathanwei05/cs2881r-dobby-qwen2.5-3b-a3-sft-3ep (revision 898bdaa987f47fcc49d0ee6171a2d5c42db54c63) against a DeepSeek V4.1 Flash judge.

Reward = persona x quality / 16, each 0-4, math correctness deliberately excluded. Selected on 100 held-out dev prompts by mean judge reward (earlier checkpoint wins ties); step 25 of 100 was selected.

Held-out results (paired, identical prompts, greedy, 2048-token cap)

SetNPersona /4 (SFT to RLAIF)Quality /4AccuracyReward delta [95% CI]
chat2002.805 -> 2.9352.220 -> 2.285-+0.026 [-0.011, +0.063]
gsm8k13193.658 -> 3.8663.332 -> 3.4540.707 -> 0.704+0.074 [+0.062, +0.087]
math5005003.816 -> 3.8402.690 -> 2.7980.384 -> 0.392+0.031 [+0.009, +0.054]

Persona and judged quality improve on math prompts with no change in accuracy. The mechanism is repair of persona dropout (persona=0 fell 4.9% to 1.7% on GSM8K), not keyword stuffing. Chat did not improve reliably and its 26.5% looping rate is unchanged.

Training: 100 updates, LoRA r=16 all-linear, LR 5e-6, beta 0.04, 32 completions/update. Final KL from the SFT reference 0.0009; max relative weight delta 0.00012.

adapter/ holds the LoRA adapter; the root is the merged model.