Moonlight / Muon is Scalable
Shows Muon optimizer scales to a 16B-parameter MoE trained on 5.7T tokens, with about 2x compute efficiency over AdamW; open checkpoints and code.
Two fixes (weight decay, per-parameter update-scale matching) let Muon run at scale without hyperparameter retuning. First public large-scale Muon MoE; the groundwork for K2's MuonClip.
- Date
- Monday, 24 February 2025
- Lab
- Moonshot AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Compute efficiency vs AdamW | ~2x scaling-law experiments, compute-optimal training | company |
| Model | 3B active / 16B total MoE, 5.7T tokens Moonlight | company |
arXiv v1 2025-02-24; HF repo created 2025-02-22. Muon itself was introduced by Keller Jordan and collaborators; Moonshot's contribution is scaling it. MIT license.
Sources
This record was checked against its sources on 6 October 2026. How we check