Muon is Scalable for LLM Training (Moonlight)
Shows the Muon optimizer scales to LLMs with weight decay and per-parameter update scaling, about 2x the compute efficiency of AdamW.
Muon (orthogonalised momentum updates, introduced by Keller Jordan on small models in 2024) had only worked at small scale. Moonshot fixed its scaling issues and trained Moonlight (3B/16B MoE, 5.7T tokens). It then used MuonClip for Kimi K2; DeepSeek-V4 also uses Muon.
- Date
- Monday, 24 February 2025
- Lab
- Moonshot AI
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Compute efficiency vs AdamW | ~2x compute-optimal training | authors |
arXiv v1 2025-02-24. Muon origin credited to Keller Jordan's NanoGPT-speedrun work; his post reports CIFAR-10 3.3 to 2.6 A100-seconds and a 1.35x NanoGPT speedup. The post's date was not shown in my fetch.
Sources
This record was checked against its sources on 6 October 2026. How we check