Mixtral of Experts (paper)
Documents Mixtral 8x7B: a sparse MoE with 47B total and 13B active parameters that matches or beats Llama 2 70B and GPT-3.5.
First strong open-weights MoE, under Apache 2.0. Top-2 of 8 experts per token per layer gives 70B-class quality at roughly 13B-class inference cost, reviving MoE for the open community and prefiguring DeepSeekMoE, Qwen MoE and Llama 4.
- Date
- Monday, 8 January 2024
- Lab
- Mistral AI
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters (total / active per token) | 47B / 13B 32K context | company |
arXiv v1 2024-01-08; the weights were already released in December 2023, so the paper documents rather than debuts the model. Authors incl. Albert Jiang, Arthur Mensch.
Sources
This record was checked against its sources on 6 October 2026. How we check