CoolFace
Modelpublic

ryankim17920/qwen3p5-2b-luspo-p25-adamw-lr1e6-a05

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes12downloads
Model Card

Qwen3.5-2B - luspo/p25 (adamw)

vs base Qwen3.5-2B: InD acc 83.8→77.1, total output tokens 3240→343 (-89%)

gpqa_diamond (OOD) acc 7.6→21.2 (+13.6 pp, +180%)

Trained via GRPO with luspo loss, p25 reward shape (alpha=0.05), adamw optimizer, lr=1.0e-06, G=8, maxsteps=200, maxcompletion_length=8000, evaluated over 3 seeds.

Accuracy vs base Qwen3.5-2B

DatasetBaseTuned (mean ± std)Δ (pp, rel %)
gsm8k81.252.0 ± 3.6-29.2 pp, -36%
arc_challenge85.279.7 ± 2.0-5.5 pp, -6%
arc_easy97.793.8 ± 1.0-3.8 pp, -4%
commonsenseqa69.870.8 ± 2.5+1.0 pp, +1%
openbookqa82.374.7 ± 2.5-7.7 pp, -9%
qasc76.575.8 ± 1.5-0.7 pp, -1%
sciq93.892.8 ± 0.3-1.0 pp, -1%
mmlu_pro(OOD)33.834.7 ± 0.3+0.8 pp, +2%
mmlu_redux(OOD)52.549.2 ± 3.3-3.3 pp, -6%
gpqa_diamond(OOD)7.621.2 ± 3.3+13.6 pp, +180%
InD Average83.877.1 ± 0.4-6.7 pp, -8%
OOD31.435.1 ± 2.3+3.7 pp, +12%
ALL68.164.5 ± 0.9-3.6 pp, -5%

Δ shows the absolute change in accuracy points (`pp`) and the relative percent change `(tuned − base) / base × 100` (`rel %`, shown as `n/a` when base accuracy is 0).

Output tokens (total) vs base Qwen3.5-2B

DatasetBaseTuned (mean ± std)Reduction %
gsm8k4450486 ± 84-89%
arc_challenge3157319 ± 13-90%
arc_easy1871291 ± 11-84%
commonsenseqa3949361 ± 10-91%
openbookqa3378317 ± 55-91%
qasc3932373 ± 14-91%
sciq1944256 ± 20-87%
mmlu_pro(OOD)65821836 ± 151-72%
mmlu_redux(OOD)55891514 ± 362-73%
gpqa_diamond(OOD)80012862 ± 487-64%
InD Average3240343 ± 3-89%
OOD67202068 ± 255-69%
ALL4281859 ± 77-80%

Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. `+`, means the tuned model generates more tokens).