AI Research Atlas

Moonlight / Muon is Scalable

Moonshot AI · 24 February 2025

Shows Muon optimizer scales to a 16B-parameter MoE trained on 5.7T tokens, with about 2x compute efficiency over AdamW; open checkpoints and code.

Two fixes (weight decay, per-parameter update-scale matching) let Muon run at scale without hyperparameter retuning. First public large-scale Muon MoE; the groundwork for K2's MuonClip.

Date
Monday, 24 February 2025
Lab
Moonshot AI
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Compute efficiency vs AdamW~2x
scaling-law experiments, compute-optimal training
company
Model3B active / 16B total MoE, 5.7T tokens
Moonlight
company

arXiv v1 2025-02-24; HF repo created 2025-02-22. Muon itself was introduced by Keller Jordan and collaborators; Moonshot's contribution is scaling it. MIT license.

Sources

  1. arxiv.org/abs/2502.16982
  2. huggingface.co/moonshotai/Moonlight-16B-A3B

This record was checked against its sources on 6 October 2026. How we check