XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
4295.7k
MiMo-V2.6-Distill-Qwen-9B
MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data. It covers coding, general-purpose agent tasks, visual coding, and cybersecurity. We release this SFT checkpoint as a starting point for open research in agentic reinforcement learning.
Evaluation
Results for the released SFT checkpoint, as reported in the MiMo-V2.6 technical report.
† Internal evaluation sets.
Training Data
The weighted SFT data mixture contains 77.4B total tokens, including 27.2B loss-bearing tokens.
Quickstart
For text generation, use a recent SGLang build with Qwen3.5 support. The checkpoint includes its tokenizer and MiMo v2.6 chat template.
sglang serve \
--model-path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B \
--reasoning-parser mimo \
--host 0.0.0.0 \
--port 30000Query the endpoint with thinking explicitly enabled:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:30000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B",
messages=[
{"role": "user", "content": "What is 15% of 240?"}
],
max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")Citation
@misc{mimo2026v26,
title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}