AI Research Atlas

MiniMax-Text-01 and MiniMax-VL-01

MiniMax · 15 January 2025

456B-total (45.9B active) open MoE with lightning attention in most layers; trained at 1M tokens and extrapolating to 4M at inference.

First open model at this scale to use a hybrid of linear (lightning) and softmax attention, with 32 experts. MiniMax claims parity with GPT-4o and Claude 3.5 Sonnet at a 20-32x longer context; the VL variant adds 512B vision-language training tokens.

Date
Wednesday, 15 January 2025
Lab
MiniMax
Kind
open-weights
Access
open weights (restricted license)

Figures

MeasureValueMeasured by
Total / active parameters456B / 45.9B
32 experts
company
RULER at 1M tokens0.910company
LongBench v2 (with CoT)56.5
vs GPT-4o 51.4, Claude 3.5 Sonnet 46.7
company
MMLU88.5
vs GPT-4o 85.7, Claude 3.5 Sonnet 88.3
company

Paper v1 2025-01-14; MiniMax's release notes list 2025-01-15. Benchmark comparisons are company-run. Weights under a custom model agreement (code MIT). A separate paper record exists in papers.json.

Sources

  1. arxiv.org/abs/2501.08313
  2. huggingface.co/MiniMaxAI/MiniMax-Text-01
  3. platform.minimax.io/docs/release-notes/models

This record was checked against its sources on 6 October 2026. How we check