MiMo-V2-Flash
309B-total, 15B-active open MoE that Xiaomi says rivals DeepSeek-V3.2 and Kimi-K2 at half and a third of their size; 73.4 SWE-bench Verified.
Hybrid attention interleaves 128-token sliding windows with global layers at 5:1; three multi-token-prediction layers give 2.6x decoding speedup; Multi-Teacher On-Policy Distillation merges RL-trained domain teachers into one student. Trained on 27T tokens, 256K context, MIT.
- Date
- Tuesday, 16 December 2025
- Lab
- Xiaomi (MiMo)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Verified | 73.4 | company |
| AIME 2025 | 94.1 | company |
| Decoding speedup from MTP | 2.6x 3 MTP layers | company |
| Pretraining tokens | 27T | company |
HF weights 2025-12-16 UTC (Wikipedia says 12-17); tech report arXiv 2026-01-06. Benchmarks are Xiaomi-run. A 2026-02-04 update (mimo-v2-flash-0204) lifted SWE-bench Verified to 78.6 per Xiaomi's release note.
Sources
- arxiv.org/abs/2601.02780
- huggingface.co/XiaomiMiMo/MiMo-V2-Flash
- mimo.mi.com/docs/en-US/news/previous-news/news20260212
This record was checked against its sources on 6 October 2026. How we check