cs.LGMay 26, 2026

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

Authors: Jiacheng LiJianchao TanHongtao XuJiaqi ZhangYifan LuYerui SunYuchen XieXunliang Cai

Abstract

The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term introduces curvature-sensitive corrections while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.

Explore similar work

CardsList
  1. AMO: Adaptive Muon Orthogonalization

    May 18, 2026Xinlin Zhuang, Panyi Ouyang, Yichen Li +7Muon OptimizerAdam

  2. Muown: Row-Norm Control for Muon Optimization

    May 11, 2026Kai Lion, Florian Hübler, Bingcong Li +2Muon OptimizerMuon

  3. Why Muon Outperforms Adam: A Curvature Perspective

    Jun 3, 2026Shuche Wang, Fengzhuo Zhang, Jiaxiang Li +2AdamMuon