Auxiliary-loss-free load balancing for MoE
Balances MoE expert load with a per-expert routing bias updated from recent load, avoiding the interference gradients of an auxiliary loss.
Before top-K routing, each expert's score gets a bias that rises when it is underused and falls when overloaded, with no loss term. Tested on models up to 3B parameters and 200B tokens; DeepSeek-V3 later adopted the idea at 671B scale.
- Date
- Wednesday, 28 August 2024
- Lab
- DeepSeek
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Experiment scale | up to 3B params, 200B tokens Better performance and balance than auxiliary-loss baselines, per the authors | company |
arXiv 2408.15664 (2024-08-28), authors Lean Wang, Huazuo Gao, Chenggang Zhao, Xu Sun, Damai Dai. The V3 report says it pioneers an auxiliary-loss-free strategy; later work (Qwen's global-batch load balancing, 2025-01) attacked the same problem differently.
Sources
This record was checked against its sources on 6 October 2026. How we check