Mixtral of Experts (paper)
The Mixtral paper documents the 8-expert top-2 sparse MoE: 47B total, 13B active parameters, 32K context, and claims parity with Llama 2 70B and GPT-3.5.
Public technical description of Mixtral 8x7B (26 authors), including the instruction-tuned variant that the paper says surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro and Llama 2 70B chat on human benchmarks.
- Date
- Monday, 8 January 2024
- Lab
- Mistral AI
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Active / total parameters | 13B / 47B abstract | company |
arXiv v1 date. Claims in the abstract are authors' own.
Sources
This record was checked against its sources on 6 October 2026. How we check