What Is Xiaomi MiMo-V2.6? Top Open-Weight Model (Sep 2026)
The short answer
MiMo-V2.6 is Xiaomi’s September 21, 2026 open-weight model family, and MiMo-V2.6-Pro is now the strongest open-weight model in the world by Artificial Analysis’s Intelligence Index, scoring 46, level with Grok 4.7 and ahead of every DeepSeek, Kimi and Qwen release. Both Pro (1.02T MoE, 42B active) and Flash (310B MoE, 15B active) are MIT-licensed, natively multimodal, have 1M-token context, and cost $0.435/$0.87 and $0.14/$0.28 per million tokens respectively on Xiaomi’s API. The story behind the numbers is a six-day, $3.5 million reinforcement-learning run that Xiaomi documented publicly and released as environments and code.
At a glance
| MiMo-V2.6-Pro | MiMo-V2.6-Flash | MiMo-V2.6-Distill-Qwen-9B | |
|---|---|---|---|
| Released | September 21, 2026 | September 21, 2026 | September 21, 2026 |
| Architecture | MoE, 1.02T total / 42B active | MoE, 310B total / 15B active | Dense 9B (from Qwen3.5-9B) |
| License | MIT | MIT | MIT |
| Modalities | Text, image, audio, video in | Same | Text |
| Context / max output | 1M / 128K | 1M / 128K | — |
| API price (in / out per MTok) | $0.435 / $0.87 | $0.14 / $0.28 | Self-host |
| Cache-hit input | $0.0036 | $0.0028 | — |
| Speed variant | Pro-UltraSpeed, up to 20x | — | — |
| AA Intelligence Index | 46 | Not yet published | — |
| AA cost per index task | ~$0.13 | — | — |
| Output speed (AA) | ~134 tok/s | — | — |
| RL training cost | ~$2.62M | ~$850K | — |
| API model name | mimo-v2.6-pro | mimo-v2.6-flash | — |
Why it ranks where it ranks
Artificial Analysis’s current index scale (the same one on which Claude Fable 5.1 and GPT-6 Astra score 53 and Claude Opus 5.5 scores 58):
| Model | AA Index | Weights |
|---|---|---|
| Claude Opus 5.5 | 58 | Closed |
| Claude Fable 5.1 / GPT-6 Astra | 53 | Closed |
| GPT-6 Sol | 48 | Closed |
| MiMo-V2.6-Pro | 46 | Open (MIT) |
| Grok 4.7 | 46 | Closed |
| Grok 4.6 | 44 | Closed |
| Gemini 3.8 Flash | 41 | Closed |
| DeepSeek V4.1 Flash | 39 | Open |
| DeepSeek V4.1 Pro | 36 | Open |
Xiaomi’s own framing is careful: Pro is “on par with Claude Opus 5 and GPT-5.6 Sol” on most agent benchmarks, and “there is still a gap” to Fable 5.1 and GPT-6 Astra. At $0.13 per index task, though, it sits on the intelligence-versus-cost Pareto frontier; Xiaomi claims 1/20 to 1/60 the price of overseas models at equal intelligence.
Do not mix this index version with the older 65.6-point scale used in August; scores are not comparable across versions.
The RL run is the actual news
MiMo-V2.6 is a post-training story. Starting from V2.5-class bases, Xiaomi ran 30 large reinforcement-learning steps per model in under six days, about 750,000 trajectories in total, at roughly $2.62 million for Pro and $850,000 for Flash. Each step began with 1,568 prompts, generated 16 candidate trajectories each (about 25,000 rollouts), and consumed 2.7–3.7 billion training tokens; trajectories averaged 110,000–150,000 tokens, meaning Xiaomi was reinforcing entire agent workflows, not short answers. Only 43.5% of Pro’s RL budget went to weight updates; 43.8% went to producing rollouts and 12.7% to grading them.
The measured effect on the out-of-sample DeepSWE v1.1 benchmark: Flash rose ~17 points (48.8 → 65.7) and Pro ~14 points (58.4 → 72.6). Xiaomi calls the approach “You Only RL Once”: a single mixed run across coding, general professional work, visual tasks and cybersecurity, using multiple lightweight “mini-harnesses” so the model does not overfit to one scaffold. Engineering details it published: fully asynchronous GRPO with partial rollouts, a frozen MoE router to stop expert-load drift, and a reward-hacking defence covering reward design, adversarial evaluation, anomaly detection and cross-validator checks.
Fuli Luo, the former DeepSeek researcher who leads the MiMo team, described it as likely one of the largest single RL runs by an open-source team.
What was open-sourced
- Pro and Flash weights plus the technical report on Hugging Face (
XiaomiMiMo/mimo-v26collection). - MiMo-V2.6-Distill-Qwen-9B, the local-friendly checkpoint.
- 7,000+ RL task environments across software engineering, vulnerability reproduction, knowledge work and web design/development, plus the RL training code. Applying them to the 9B distill lifted SWE-bench Verified from 61.1 to 66.2, Terminal Bench 2.1 from 37.1 to 52.8 and MiMo Cyber Bench from 31.3 to 47.0.
This is more than DeepSeek, Moonshot or Alibaba typically release alongside weights; see fully open vs open-weight models for why the environments matter.
Capabilities Xiaomi is pushing
Beyond terminal coding, V2.6 folds in 3D spatial reasoning, computer-use actions and multimodal perception: building a runnable 3D game world from a prompt via multi-agent collaboration, Blender modelling from text or reference images, controlling a Franka Panda robot arm from multi-view camera input, front-end and slide generation with Figma, and a formal Lean 4 proof of Li and Yorke’s “Period Three Implies Chaos” theorem assembled with sub-agents. It also ships alongside MiMo Desktop, which uses the UltraSpeed mode.
Trade-offs before you adopt it
- Hosting: Xiaomi’s API serves from China, which many enterprises cannot use. Self-hosting Pro means a multi-GPU node for a 1.02T-parameter MoE; Flash at 310B is still a server-class deployment. Expect US inference providers to host both within days, as they did for DeepSeek V4 and Kimi K3.
- Index gap to the closed frontier: 46 versus 53–58. For open-ended agentic work, Claude Opus 5.5 at $4/$20 is roughly 10x the price and 12 points ahead.
- Flash versus GPT-6 Luna: OpenAI’s new $0.10/$0.50 Luna undercuts MiMo-V2.6-Flash on input and is closed but US-hosted; Flash wins on output price ($0.28 vs $0.50), open weights and 1M multimodal context. See GPT-6 Sol vs Luna vs Astra.
- Distillation politics: Anthropic’s September 2026 threat report and the CISA AA26-251A advisory put Chinese open-weight labs under scrutiny; Xiaomi’s published RL provenance is partly a response to that.
Last verified: September 23, 2026.