CoolFace
Modelpublic

logan7000/mllm-mmr1-ttrl-gemma3-12b-endpoint

sourceHugging Facegemmaupdated 2mo agoView on Hugging Face
0likes7downloads
Model Card

CoRL · gemma-3-12b-it · TTRL · mmr1 (endpoint)

  • —Base: google/gemma-3-12b-it
  • —Method: TTRL — self-labeling GRPO (majority-vote pseudo-labels, no GT)
  • —Dataset: mmr1 (~8k), 1 epoch = 722 steps
  • —This checkpoint: near-endpoint, step 720 (exact endpoint step 722 was lost with the local disk; this is the last archived checkpoint, 2 steps before the true end — the endpoint eval numbers below are from the true step-722 checkpoint)
  • —Paper cell: MLLM main table (mmr1, big tier) · gemma-3-12b-it · TTRL
  • —WandB run: https://wandb.ai/logan-yang2002-johns-hopkins-university/co-rl-mllm-mmr1/runs/gemma312bttrl_mmr1-v1
  • —best repo: q1716523669/mllm-mmr1-ttrl-gemma3-12b · endpoint repo: q1716523669/mllm-mmr1-ttrl-gemma3-12b-endpoint

Eval — 4 benchmarks (prompt=answer · greedy T=0 · mathruler)

ckptMathVisionMathVerseMathVistaWe-Math**avg**
best (step 20)28.3637.6137.862.9341.68
endpoint (step 722)14.6718.9332.245.2927.77

best → endpoint: 41.68 → 27.77 — TTRL belief-reinforcement collapse (self-reward saturates while real accuracy degrades).