CoolFace
Modelpublic

ryankim17920/qwen3p5-2b-luspo-diffscaled-adamw-lr2e6

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes10downloads
Model Card

Qwen3.5-2B - luspo/difficulty_scaled (adamw)

vs base Qwen3.5-2B: InD acc 83.8→0.2, total output tokens 3240→6665 (+106%)

gpqa_diamond (OOD) acc 7.6→0.0 (-7.6 pp, -100%)

Trained via GRPO with luspo loss, difficultyscaled reward shape (alpha=0.1), adamw optimizer, lr=2.0e-06, G=8, maxsteps=200, maxcompletionlength=8000, evaluated over 3 seeds.

Accuracy vs base Qwen3.5-2B

DatasetBaseTuned (mean ± std)Δ (pp, rel %)
gsm8k81.21.7 ± 1.0-79.5 pp, -98%
arc_challenge85.20.0 ± 0.0-85.2 pp, -100%
arc_easy97.70.0 ± 0.0-97.7 pp, -100%
commonsenseqa69.80.0 ± 0.0-69.8 pp, -100%
openbookqa82.30.0 ± 0.0-82.3 pp, -100%
qasc76.50.0 ± 0.0-76.5 pp, -100%
sciq93.80.0 ± 0.0-93.8 pp, -100%
mmlu_pro(OOD)33.80.0 ± 0.0-33.8 pp, -100%
mmlu_redux(OOD)52.50.0 ± 0.0-52.5 pp, -100%
gpqa_diamond(OOD)7.60.0 ± 0.0-7.6 pp, -100%
InD Average83.80.2 ± 0.1-83.5 pp, -100%
OOD31.40.0 ± 0.0-31.4 pp, -100%
ALL68.10.2 ± 0.1-67.9 pp, -100%

Δ shows the absolute change in accuracy points (`pp`) and the relative percent change `(tuned − base) / base × 100` (`rel %`, shown as `n/a` when base accuracy is 0).

Output tokens (total) vs base Qwen3.5-2B

DatasetBaseTuned (mean ± std)Reduction %
gsm8k44505647 ± 76+27%
arc_challenge31576948 ± 229+120%
arc_easy18716641 ± 211+255%
commonsenseqa39497394 ± 91+87%
openbookqa33786480 ± 339+92%
qasc39327137 ± 183+82%
sciq19446408 ± 71+230%
mmlu_pro(OOD)65826786 ± 236+3%
mmlu_redux(OOD)55896840 ± 144+22%
gpqa_diamond(OOD)80016547 ± 193-18%
InD Average32406665 ± 100+106%
OOD67206725 ± 189+0%
ALL42816683 ± 120+56%

Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. `+`, means the tuned model generates more tokens).