AI Research Atlas

DeepSeekMoE

DeepSeek · 11 January 2024

DeepSeekMoE uses fine-grained expert segmentation plus always-on shared experts, the MoE design that V2, V3, R1 and V4 all inherit.

Splits each expert into smaller units and activates more of them, and reserves shared experts for common knowledge. The 16B model approached LLaMA2 7B at about 40% of the compute; the 145B-scale model matched DeepSeek 67B at about 28.5%.

Date
Thursday, 11 January 2024
Lab
DeepSeek
Kind
paper
Access
open weights (restricted license)

Figures

MeasureValueMeasured by
Compute vs dense peer (16B)~40%
Approaches LLaMA2 7B quality
company
Compute vs dense peer (145B)~28.5%
Comparable to DeepSeek 67B
company

DeepSeekMoE-16B base weights appeared on Hugging Face 2024-01-08 (repo creation date), a few days before the arXiv paper.

Sources

  1. arxiv.org/abs/2401.06066
  2. huggingface.co/api/models/deepseek-ai/deepseek-moe-16b-base

This record was checked against its sources on 6 October 2026. How we check