Mixtral 8x7B
Mixtral 8x7B, a sparse mixture-of-experts with 12.9B active of 46.7B parameters, matches or beats Llama 2 70B and GPT-3.5 under Apache 2.0.
A router picks 2 of 8 expert blocks per token, giving 70B-class quality at about 13B-class inference cost with 6x faster inference than Llama 2 70B. 32K context. It brought sparse MoE into everyday open-weights practice.
- Date
- Monday, 11 December 2023
- Lab
- Mistral AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Active / total parameters | 12.9B / 46.7B 2 of 8 experts per token | company |
| MT-Bench (instruct) | 8.30 | company |
Paper followed 2024-01-08 (arXiv 2401.04088).
Sources
This record was checked against its sources on 6 October 2026. How we check