MiniMax-Text-01 and MiniMax-VL-01
456B-total (45.9B active) open MoE with lightning attention in most layers; trained at 1M tokens and extrapolating to 4M at inference.
First open model at this scale to use a hybrid of linear (lightning) and softmax attention, with 32 experts. MiniMax claims parity with GPT-4o and Claude 3.5 Sonnet at a 20-32x longer context; the VL variant adds 512B vision-language training tokens.
- Date
- Wednesday, 15 January 2025
- Lab
- MiniMax
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Total / active parameters | 456B / 45.9B 32 experts | company |
| RULER at 1M tokens | 0.910 | company |
| LongBench v2 (with CoT) | 56.5 vs GPT-4o 51.4, Claude 3.5 Sonnet 46.7 | company |
| MMLU | 88.5 vs GPT-4o 85.7, Claude 3.5 Sonnet 88.3 | company |
Paper v1 2025-01-14; MiniMax's release notes list 2025-01-15. Benchmark comparisons are company-run. Weights under a custom model agreement (code MIT). A separate paper record exists in papers.json.
Sources
- arxiv.org/abs/2501.08313
- huggingface.co/MiniMaxAI/MiniMax-Text-01
- platform.minimax.io/docs/release-notes/models
This record was checked against its sources on 6 October 2026. How we check