HumorR1/policy-qwen3vl-2b-grpo-newyorker
04
Add 8192-token eval JSON
Update model card with honest 8192-token eval (was clipped at 2048)
Add eval JSON for the SFT+GRPO policy
Replace with SFT-think + GRPO policy (+2.30σ over base, restores thinking format)
Add 50×5 held-out eval (RM scores per cartoon × sample)
GRPO policy (Qwen3-VL-2B-Thinking + LoRA, FB=0.1, 200 steps)
initial commit
