Mixtral 8x22B
Mixtral 8x22B activates 39B of 141B parameters with 64K context under Apache 2.0, scoring 90.8% GSM8K and 44.6% MATH in instruct form.
A larger sparse MoE with native function calling, 64K context and stronger math and code than Mixtral 8x7B, under Apache 2.0.
- Date
- Wednesday, 17 April 2024
- Lab
- Mistral AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| GSM8K (instruct) | 90.8% | company |
| MATH (instruct) | 44.6% | company |
Sources
This record was checked against its sources on 6 October 2026. How we check