ryankim17920/qwen3p5-2b-luspo-p25-adamw-lr1e6-a05
012
Qwen3.5-2B - luspo/p25 (adamw)
vs base Qwen3.5-2B: InD acc 83.8→77.1, total output tokens 3240→343 (-89%)
gpqa_diamond (OOD) acc 7.6→21.2 (+13.6 pp, +180%)
Trained via GRPO with luspo loss, p25 reward shape (alpha=0.05), adamw optimizer, lr=1.0e-06, G=8, maxsteps=200, maxcompletion_length=8000, evaluated over 3 seeds.
Accuracy vs base Qwen3.5-2B
Δ shows the absolute change in accuracy points (`pp`) and the relative percent change `(tuned − base) / base × 100` (`rel %`, shown as `n/a` when base accuracy is 0).
Output tokens (total) vs base Qwen3.5-2B
Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. `+`, means the tuned model generates more tokens).
