DeepSeekMoE
DeepSeekMoE uses fine-grained expert segmentation plus always-on shared experts, the MoE design that V2, V3, R1 and V4 all inherit.
Splits each expert into smaller units and activates more of them, and reserves shared experts for common knowledge. The 16B model approached LLaMA2 7B at about 40% of the compute; the 145B-scale model matched DeepSeek 67B at about 28.5%.
- Date
- Thursday, 11 January 2024
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Compute vs dense peer (16B) | ~40% Approaches LLaMA2 7B quality | company |
| Compute vs dense peer (145B) | ~28.5% Comparable to DeepSeek 67B | company |
DeepSeekMoE-16B base weights appeared on Hugging Face 2024-01-08 (repo creation date), a few days before the arXiv paper.
Sources
This record was checked against its sources on 6 October 2026. How we check