ryankim17920/qwen3p5-2b-luspo-diffscaled-adamw-lr2e6
010
Qwen3.5-2B - luspo/difficulty_scaled (adamw)
vs base Qwen3.5-2B: InD acc 83.8→0.2, total output tokens 3240→6665 (+106%)
gpqa_diamond (OOD) acc 7.6→0.0 (-7.6 pp, -100%)
Trained via GRPO with luspo loss, difficultyscaled reward shape (alpha=0.1), adamw optimizer, lr=2.0e-06, G=8, maxsteps=200, maxcompletionlength=8000, evaluated over 3 seeds.
Accuracy vs base Qwen3.5-2B
Δ shows the absolute change in accuracy points (`pp`) and the relative percent change `(tuned − base) / base × 100` (`rel %`, shown as `n/a` when base accuracy is 0).
Output tokens (total) vs base Qwen3.5-2B
Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. `+`, means the tuned model generates more tokens).
