MiniMax M3
Natively multimodal ~428B-total (~23B active) model with MiniMax Sparse Attention for 1M context at about 1/20 the per-token cost of M2; 80.5% SWE-bench Verified.
MSA selects key-value blocks instead of full attention, which gives 9x prefill and 15x decode speedups versus M2 at 1M tokens. A separate arXiv paper (2606.13392) credits MSA with a 28.4x attention-compute cut. Weights released in June under a community licence.
- Date
- Monday, 1 June 2026
- Lab
- MiniMax
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Verified | 80.5% | company |
| SWE-bench Pro | 59.0% | company |
| Terminal-Bench 2.1 | 66.0% | company |
| MMMU Pro | 78.1% | company |
| Prefill / decode speedup vs M2 at 1M context | 9x / 15x | company |
API launch 2026-06-01; weights and technical report followed within about 10 days (paper v1 2026-06-11). Parameter counts are from the HF card; the blog gives none. The paper's 109B figure refers to a research model for MSA, not M3. Benchmarks are company-run.
Sources
This record was checked against its sources on 6 October 2026. How we check