MiniMax-M1
Open 456B hybrid-attention reasoning model with 1M-token context and the CISPO RL algorithm; full RL run cost a reported $534,700.
Scales the Text-01 architecture with large-scale RL. CISPO clips importance-sampling weights rather than token updates; 512 H800 GPUs finished RL in three weeks. At 100K generated tokens it uses 25% of the FLOPs of DeepSeek-R1. Apache 2.0.
- Date
- Monday, 16 June 2025
- Lab
- MiniMax
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| RL training cost | $534,700 512 H800 GPUs, 3 weeks | company |
| SWE-bench Verified (M1-80K) | 56.0% | company |
| AIME 2024 (M1-80K) | 86.0% | company |
| OpenAI-MRCR (1M tokens) | 56.2% | company |
| FLOPs vs DeepSeek-R1 at 100K tokens | 25% | company |
Paper v1 2025-06-16; trade tracker lists 2025-06-17. Two checkpoints with 40K and 80K thinking budgets. Efficiency and benchmark numbers are company-reported.
Sources
This record was checked against its sources on 6 October 2026. How we check