AI Research Atlas

MiMo-V2-Flash

Xiaomi (MiMo) · 16 December 2025

309B-total, 15B-active open MoE that Xiaomi says rivals DeepSeek-V3.2 and Kimi-K2 at half and a third of their size; 73.4 SWE-bench Verified.

Hybrid attention interleaves 128-token sliding windows with global layers at 5:1; three multi-token-prediction layers give 2.6x decoding speedup; Multi-Teacher On-Policy Distillation merges RL-trained domain teachers into one student. Trained on 27T tokens, 256K context, MIT.

Date
Tuesday, 16 December 2025
Lab
Xiaomi (MiMo)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
SWE-bench Verified73.4company
AIME 202594.1company
Decoding speedup from MTP2.6x
3 MTP layers
company
Pretraining tokens27Tcompany

HF weights 2025-12-16 UTC (Wikipedia says 12-17); tech report arXiv 2026-01-06. Benchmarks are Xiaomi-run. A 2026-02-04 update (mimo-v2-flash-0204) lifted SWE-bench Verified to 78.6 per Xiaomi's release note.

Sources

  1. arxiv.org/abs/2601.02780
  2. huggingface.co/XiaomiMiMo/MiMo-V2-Flash
  3. mimo.mi.com/docs/en-US/news/previous-news/news20260212

This record was checked against its sources on 6 October 2026. How we check