axel-sdq/cs2881r-dobby-qwen2.5-3b-rlaif-grpo100-step25
Dobby Qwen2.5-3B — RLAIF (GRPO) checkpoint 25
Harvard CS 2881R Assignment 1, checkpoint 2. LoRA GRPO from nathanwei05/cs2881r-dobby-qwen2.5-3b-a3-sft-3ep (revision 898bdaa987f47fcc49d0ee6171a2d5c42db54c63) against a DeepSeek V4.1 Flash judge.
Reward = persona x quality / 16, each 0-4, math correctness deliberately excluded. Selected on 100 held-out dev prompts by mean judge reward (earlier checkpoint wins ties); step 25 of 100 was selected.
Held-out results (paired, identical prompts, greedy, 2048-token cap)
Persona and judged quality improve on math prompts with no change in accuracy. The mechanism is repair of persona dropout (persona=0 fell 4.9% to 1.7% on GSM8K), not keyword stuffing. Chat did not improve reliably and its 26.5% looping rate is unchanged.
Training: 100 updates, LoRA r=16 all-linear, LR 5e-6, beta 0.04, 32 completions/update. Final KL from the SFT reference 0.0009; max relative weight delta 0.00012.
adapter/ holds the LoRA adapter; the root is the merged model.
