MiniMax-01 (lightning attention)
456B MoE (45.9B active) with lightning linear attention in most layers, trained at 1M tokens and extrapolating to 4M, open-sourced.
First open frontier-scale model built around linear attention (mostly linear, a softmax layer every eighth). Claimed GPT-4o/Claude-3.5-class quality at 20-32x the context. Notably MiniMax later reverted to full attention for M2, an honest data point on hybrid limits.
- Date
- Tuesday, 14 January 2025
- Lab
- MiniMax
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters (total / active) | 456B / 45.9B 32 experts; 1M training context, 4M at inference | company |
arXiv v1 2025-01-14. Code and models open-sourced per the abstract. MiniMax's later M2 series is a separate record in the models dataset; whether it reverted to full attention was not verified here.
Sources
This record was checked against its sources on 6 October 2026. How we check