AI Research Atlas

Mixtral of Experts (paper)

Mistral AI · 8 January 2024

Documents Mixtral 8x7B: a sparse MoE with 47B total and 13B active parameters that matches or beats Llama 2 70B and GPT-3.5.

First strong open-weights MoE, under Apache 2.0. Top-2 of 8 experts per token per layer gives 70B-class quality at roughly 13B-class inference cost, reviving MoE for the open community and prefiguring DeepSeekMoE, Qwen MoE and Llama 4.

Date
Monday, 8 January 2024
Lab
Mistral AI
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
Parameters (total / active per token)47B / 13B
32K context
company

arXiv v1 2024-01-08; the weights were already released in December 2023, so the paper documents rather than debuts the model. Authors incl. Albert Jiang, Arthur Mensch.

Sources

  1. arxiv.org/abs/2401.04088

This record was checked against its sources on 6 October 2026. How we check