MiMo-7B
Xiaomi's first reasoning LLM: a 7B model pretrained for reasoning (25T tokens) and RL-tuned on 130K verifiable problems, claimed to beat OpenAI o1-mini.
Pretraining adds multi-token prediction and reasoning-dense data; RL uses a difficulty-based reward to handle sparse rewards. The base model is said to beat 32B models. MIT licence; paper arXiv 2505.07608 (2025-05-12).
- Date
- Wednesday, 30 April 2025
- Lab
- Xiaomi (MiMo)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Pretraining tokens | 25T | company |
| Verifiable RL problems | 130K math and programming | company |
Release 2025-04-30 per Wikipedia (HF repo 2025-04-29 UTC). Claims vs o1-mini are authors' own. Team led by Luo Fuli per secondary reporting (not verified here).
Sources
This record was checked against its sources on 6 October 2026. How we check