Mistral Small 4
Mistral Small 4, a 119B MoE with 6B active parameters, merges reasoning, vision and agentic coding into one Apache 2.0 model.
Folds Magistral, Pixtral and Devstral into a single model with configurable reasoning effort, 256K context and 128 experts (4 active). Mistral claims 40% lower latency and 3x throughput versus Small 3, and parity with GPT-OSS 120B on shorter outputs.
- Date
- Monday, 16 March 2026
- Lab
- Mistral AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Active / total parameters | 6B / 119B 128 experts, 4 active per token | company |
| Context window | 256K tokens | company |
Announced alongside an NVIDIA partnership the same day.
Sources
This record was checked against its sources on 6 October 2026. How we check