math.OCMay 10, 2026

Phases of Muon: When Muon Eclipses SignSGD

Authors: Elliot PaquetteNoah MarshallLucas BenigniGuangyuan WangAtish AgarwalaCourtney Paquette

Organizations: Mathematics and Statistics Department, McGill University · Mathematics and Statistics Department, Université de Montréal · Google DeepMind

Abstract

Recently, Muon and related spectral optimizers have demonstrated strong empirical performance as scalable stochastic methods, often outperforming Adam. Yet their behaviour remains poorly understood. We analyze stochastic spectral optimizers, including Muon, on a high-dimensional matrix-valued least squares problem. We derive explicit deterministic dynamics that provide a tractable framework for studying learning behaviour with a focus on (stochastic) SignSVD, which Muon approximates, and (stochastic) SignSGD, the latter serving as a proxy for Adam. Our analysis shows that for large batch size, SignSVD performs a square-root preconditioning with respect to the data covariance spectrum, while for small batch size smaller eigenmodes behave like SGD, slowing down convergence. We contrast with SignSGD which for generic covariance performs no preconditioning and has no transition, leading to different optimal learning rates and convergence characteristics. The two methods match up to a constant factor with isotropic data, but behave differently with anisotropic data. An analysis of a power law covariance model with data exponent αα and target exponent ββ shows there are three phases in the (α,β)(α,β) plane: one where SignSGD is uniformly favored, one where SignSVD is uniformly favored, and a third where the two methods exhibit a trade-off in performance.

Explore similar work

CardsList
  1. Beyond the Matrix Sign: Quadratic Spectral Descent

    Sep 7, 2026Qiaozhe Zhang, Jun Sun, Yingzhuang LiuMuon OptimizerMuon

  2. LionMuon: Alternating Spectral and Sign Descent for Efficient Training

    May 19, 2026Arman Bolatov, Artem Riabinin, Nikita Kornilov +4Muon OptimizerMuon