AI Research Atlas

DeepSeekMoE

DeepSeek · 11 January 2024

DeepSeekMoE uses fine-grained expert segmentation plus always-on shared experts, and a 16B MoE matches Llama 2 7B at about 40% of the compute.

Splits experts into many smaller ones (activate mK of mN) and isolates shared experts for common knowledge, pushing expert specialisation. The design carried into DeepSeek-V2, V3 and R1, and influenced later open MoEs (Kimi K2, Qwen3 MoE).

Date
Thursday, 11 January 2024
Lab
DeepSeek
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
DeepSeekMoE 16B vs Llama2 7B~40% of compute
comparable performance
authors

arXiv v1 2024-01-11. Lead author Damai Dai. The 145B-scale cost ratio in the abstract is garbled in my fetch (28.5% vs 18.2%), so I did not record it.

Sources

  1. arxiv.org/abs/2401.06066

This record was checked against its sources on 6 October 2026. How we check

Related