cs.LGSep 30, 2026

Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

Authors: Hanyng Peng, Hui Wang, Yue Yu

Organizations: Pengcheng Laboratory, Shenzhen, China

Abstract

Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent (L0,L1,q)(\mathcal L_0,\mathcal L_1,q)-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an O(T−1/4)\mathcal O(T^{-1/4}) bound on the expected average Frobenius gradient norm. The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values. Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed. The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift. For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds. A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition. The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.

Figures & tables

Explore similar work

CardsList
  1. Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition

    May 7, 2026Sayantan Choudhury, Xiaoran Cheng, Martin Takáč +2Momentum Stochastic Gradient DescentMuon

  2. Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence

    May 13, 2026Tien-Phat Nguyen, Truong Nguyen, Minh-Phuc Truong +3MuonNewton

  3. OptMuon: Closed-Loop Orthogonalized Momentum Methods for Stochastic Optimization with Zero-Noise Optimality

    Jun 7, 2026Ganzhao YuanMuonMomentum Stochastic Gradient Descent