AI Research Atlas

Auxiliary-loss-free load balancing for MoE

DeepSeek · 28 August 2024

Balances MoE expert load with a per-expert routing bias updated from recent load, avoiding the interference gradients of an auxiliary loss.

Before top-K routing, each expert's score gets a bias that rises when it is underused and falls when overloaded, with no loss term. Tested on models up to 3B parameters and 200B tokens; DeepSeek-V3 later adopted the idea at 671B scale.

Date
Wednesday, 28 August 2024
Lab
DeepSeek
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
Experiment scaleup to 3B params, 200B tokens
Better performance and balance than auxiliary-loss baselines, per the authors
company

arXiv 2408.15664 (2024-08-28), authors Lean Wang, Huazuo Gao, Chenggang Zhao, Xu Sun, Damai Dai. The V3 report says it pioneers an auxiliary-loss-free strategy; later work (Qwen's global-batch load balancing, 2025-01) attacked the same problem differently.

Sources

  1. arxiv.org/abs/2408.15664
  2. arxiv.org/abs/2412.19437

This record was checked against its sources on 6 October 2026. How we check

Related