Global-batch load balancing for MoE training
Computing the MoE balance loss over the global batch, not each micro-batch, lets experts specialise and improves training with almost no cost.
Micro-batch balancing forces uniform expert use even on narrow data such as one code batch. Synchronising expert-frequency statistics across the global batch relaxes that, which Qwen3-Max later cites as part of its design.
- Date
- Tuesday, 21 January 2025
- Lab
- Alibaba (Qwen)
- Kind
- paper
- Access
- research preview
Blog dated 2025-01-21. Companion paper not opened. A parallel answer to DeepSeek's auxiliary-loss-free balancing; Qwen3-Max's blog says its architecture incorporates the global-batch loss.
Sources
This record was checked against its sources on 6 October 2026. How we check