AI Research Atlas

MiniMax M3

MiniMax · 1 June 2026

Natively multimodal ~428B-total (~23B active) model with MiniMax Sparse Attention for 1M context at about 1/20 the per-token cost of M2; 80.5% SWE-bench Verified.

MSA selects key-value blocks instead of full attention, which gives 9x prefill and 15x decode speedups versus M2 at 1M tokens. A separate arXiv paper (2606.13392) credits MSA with a 28.4x attention-compute cut. Weights released in June under a community licence.

Date
Monday, 1 June 2026
Lab
MiniMax
Kind
open-weights
Access
open weights (restricted license)

Figures

MeasureValueMeasured by
SWE-bench Verified80.5%company
SWE-bench Pro59.0%company
Terminal-Bench 2.166.0%company
MMMU Pro78.1%company
Prefill / decode speedup vs M2 at 1M context9x / 15xcompany

API launch 2026-06-01; weights and technical report followed within about 10 days (paper v1 2026-06-11). Parameter counts are from the HF card; the blog gives none. The paper's 109B figure refers to a research model for MSA, not M3. Benchmarks are company-run.

Sources

  1. www.minimax.io/blog/minimax-m3
  2. huggingface.co/MiniMaxAI/MiniMax-M3
  3. arxiv.org/abs/2606.13392

This record was checked against its sources on 6 October 2026. How we check

Related