MiMo V2 Flash
MiMo V2 Flash is Xiaomi's open-weight reasoning model delivering frontier class capability at a fraction of the inference cost of dense alternatives. The 309 billion parameter Sparse MoE architecture activates only 15 billion parameters per token, reaching up to 3 times faster generation than comparable models while staying competitive across reasoning, coding, and tool-use benchmarks. Released in December 2025, the model was pretrained on 27 trillion tokens using Multi-Token Prediction, starting from a native 32K context length later extended to 256K. Its hybrid attention mechanism interleaves sliding window and global attention at a 5:1 ratio with a 128 token window, cutting KV-cache storage by roughly 6 times while preserving long range reasoning. MiMo V2 Flash ranks among the leading open-source models on code and agent benchmarks, scoring 73.4 percent on SWE-Bench Verified and 71.7 percent on SWE-Bench Multilingual, comparable to several strong closed-source systems. It rivals top open-weight alternatives such as DeepSeek-V3.2 and Kimi-K2 while using a far smaller total parameter count. Its post-training uses a Multi-Teacher On-Policy Distillation paradigm, where specialized teachers supply token level rewards for efficient knowledge transfer. The model is open-sourced under the MIT license with weights on Hugging Face and inference code on GitHub.
Key info
Available routes
No routes currently available — MiMo V2 Flash isn't routed through the Opper gateway right now. It may return.
Contact us about this model →