AI Research Atlas

Muon is Scalable for LLM Training (Moonlight)

Moonshot AI · 24 February 2025

Shows the Muon optimizer scales to LLMs with weight decay and per-parameter update scaling, about 2x the compute efficiency of AdamW.

Muon (orthogonalised momentum updates, introduced by Keller Jordan on small models in 2024) had only worked at small scale. Moonshot fixed its scaling issues and trained Moonlight (3B/16B MoE, 5.7T tokens). It then used MuonClip for Kimi K2; DeepSeek-V4 also uses Muon.

Date
Monday, 24 February 2025
Lab
Moonshot AI
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
Compute efficiency vs AdamW~2x
compute-optimal training
authors

arXiv v1 2025-02-24. Muon origin credited to Keller Jordan's NanoGPT-speedrun work; his post reports CIFAR-10 3.3 to 2.6 A100-seconds and a 1.35x NanoGPT speedup. The post's date was not shown in my fetch.

Sources

  1. arxiv.org/abs/2502.16982
  2. kellerjordan.github.io/posts/muon/

This record was checked against its sources on 6 October 2026. How we check