AI Research Atlas

Mixtral 8x7B

Mistral AI · 11 December 2023

Mixtral 8x7B, a sparse mixture-of-experts with 12.9B active of 46.7B parameters, matches or beats Llama 2 70B and GPT-3.5 under Apache 2.0.

A router picks 2 of 8 expert blocks per token, giving 70B-class quality at about 13B-class inference cost with 6x faster inference than Llama 2 70B. 32K context. It brought sparse MoE into everyday open-weights practice.

Date
Monday, 11 December 2023
Lab
Mistral AI
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Active / total parameters12.9B / 46.7B
2 of 8 experts per token
company
MT-Bench (instruct)8.30company

Paper followed 2024-01-08 (arXiv 2401.04088).

Sources

  1. mistral.ai/news/mixtral-of-experts
  2. arxiv.org/abs/2401.04088

This record was checked against its sources on 6 October 2026. How we check

Related