logan7000/mllm-mmr1-ttrl-gemma3-12b-endpoint
07
CoRL · gemma-3-12b-it · TTRL · mmr1 (endpoint)
- Base: google/gemma-3-12b-it
- Method: TTRL — self-labeling GRPO (majority-vote pseudo-labels, no GT)
- Dataset: mmr1 (~8k), 1 epoch = 722 steps
- This checkpoint: near-endpoint, step 720 (exact endpoint step 722 was lost with the local disk; this is the last archived checkpoint, 2 steps before the true end — the endpoint eval numbers below are from the true step-722 checkpoint)
- Paper cell: MLLM main table (mmr1, big tier) · gemma-3-12b-it · TTRL
- WandB run: https://wandb.ai/logan-yang2002-johns-hopkins-university/co-rl-mllm-mmr1/runs/gemma312bttrl_mmr1-v1
- best repo:
q1716523669/mllm-mmr1-ttrl-gemma3-12b· endpoint repo:q1716523669/mllm-mmr1-ttrl-gemma3-12b-endpoint
Eval — 4 benchmarks (prompt=answer · greedy T=0 · mathruler)
best → endpoint: 41.68 → 27.77 — TTRL belief-reinforcement collapse (self-reward saturates while real accuracy degrades).
