CoolFace
Modelpublic

arianraje/mimo-7b-mamba3-hybrid-stage3-opd-1000m-bias0

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes265downloads
Model Card

mimo-7b-mamba3-hybrid-stage3-opd-1000m-bias0

Stage-3 on-policy distillation (OPD) endpoint of the MiMo-7B-RL-0530 x Mamba-3 (MIMO) hybrid at 1.0B OPD tokens (WSD ladder rung wsd_mamba3_mimo_1000m_b200). The student is the uniform-1:4 Mamba-3 conversion of MiMo-7B-RL-0530 (27 Mamba-3 MIMO mixers, rank 4, the SAME mixer class as the Qwen3-4B arm; 9 retained Qwen2-style full-attention layers with their q/k/v biases; the converted layers' q/k/v biases are dropped and the B/C norms start at 1 since MiMo has no q/k norm); the teacher is the unmodified XiaomiMiMo/MiMo-7B-RL-0530, scoring the student's own rollouts with a per-token reverse KL over the full vocabulary. The stage-2b input of this ladder is the on-disk stage-2b checkpoint mimo-7b-mamba3-u4-stage2b-v0 (stage 1 -> 2a -> 2b on the MiMo memmaps: 100M align at microbatch 4, 600M KD at 5e-5, 294M KD at 32k at 2.5e-5; not itself on the Hub). Read this arm DOWN the MiMo column (against MiMo GDN and MiMo Mamba2 at the same budget) and ACROSS the Mamba-3 row (against the Qwen3-4B Mamba-3 ladder, same mixer, other teacher). The ladder follows the MiMo column's recipe (flat LR 1e-5, maxseq 65536, frac_rug 0.20) with the Qwen ladder's sampler batch (256) from the 200M rung's second half on.

consumed OPD tokens1,000,106,637 (budget 1,000,000,000)
trainer step6092
rollout horizon at the end32768 tokens (horizon_max 32768)
LRflat 1e-05 (all params), linear decay over the last 482 x 131,072 = 63.2M tokens to 0
promptsstage3_prompts_v1 (OpenThoughts-114k-math + dolly-15k; fracgeneral 0.2, fracrug 0.2)
samplervLLM (native Mamba-3 engine), temperature 1 / top-p 1, gen_batch 32
W&Baraje/mimo-mamba3-linearization run mimo-mamba3-wsd-1000m-b200-20260916
codeLinearization branch mimo-mamba3 @ 414ba55

Closing held-out val: see stage3_wsd_summary.json.

Model weights only (bf16, snapshots/final of the run; the pre-decay full-state branch point end-minus-0482 is not uploaded). No downstream evaluation is included; evals run elsewhere.

Loading

This is a custom container (model_type: mimo_mamba3), not a stock architecture, so trust_remote_code=True is required. The modeling code is bundled in this repo (modeling_qwen3_mamba3.py, modeling_mimo_mamba3.py; auto_map points at modeling_mimo_mamba3.py) and needs no other checkout:

python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("arianraje/mimo-7b-mamba3-hybrid-stage3-opd-1000m-bias0", trust_remote_code=True, dtype="bfloat16")
tok = AutoTokenizer.from_pretrained("arianraje/mimo-7b-mamba3-hybrid-stage3-opd-1000m-bias0")

Kernel. The mixers call upstream's fused Mamba-3 MIMO kernel (mamba_ssm.ops.tilelang.mamba3.mamba3_mimo) from state-spaces/mamba at commit e9594ce (TileLang + Triton; the pip mamba_ssm 2.2.6.post3 predates Mamba-3, so install from source). Without it the model warns once and every mixer runs the sequential fp32 torch reference: numerically equivalent, slow. For serving, the Linearization repo's src/models/vllm_mimo_mamba3.py is a native vLLM 0.11.2 implementation of this architecture.