DeepSeekMoE
DeepSeekMoE uses fine-grained expert segmentation plus always-on shared experts, and a 16B MoE matches Llama 2 7B at about 40% of the compute.
Splits experts into many smaller ones (activate mK of mN) and isolates shared experts for common knowledge, pushing expert specialisation. The design carried into DeepSeek-V2, V3 and R1, and influenced later open MoEs (Kimi K2, Qwen3 MoE).
- Date
- Thursday, 11 January 2024
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| DeepSeekMoE 16B vs Llama2 7B | ~40% of compute comparable performance | authors |
arXiv v1 2024-01-11. Lead author Damai Dai. The 145B-scale cost ratio in the abstract is garbled in my fetch (28.5% vs 18.2%), so I did not record it.
Sources
This record was checked against its sources on 6 October 2026. How we check