Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum
Organizations: Pengcheng Laboratory, Shenzhen, China
Abstract
Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent -smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an bound on the expected average Frobenius gradient norm. The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values. Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed. The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift. For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds. A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition. The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.
Figures & tables
| Work | Block structure | Orthogonalization in theory | Nesterov | Smoothness and noise | Stationarity and dimension factors |
|---|---|---|---|---|---|
| Pethick et al. (2025) (Scion/uSCG) | General-norm theorem; layer-wise design | Exact LMO; polar factor for Muon | No a | Norm smoothness; bounded Euclidean variance | Dual-norm gradient; geometry-dependent constants |
| Riabinin et al. (2025) (Gluon) | Explicitly layer-wise | Exact block LMOs | No | Block generalized smoothness; block dual-norm variance | Weighted block dual norms; norm-specific constants |
| Shen et al. (2025) | Single matrix | Exact SVD-polar factor | No | Frobenius or spectral smoothness; bounded Frobenius variance | Frobenius/nuclear gradient; rank factors in stochastic bounds |
| Kim and Oh (2025) | Single matrix | Finite Taylor NS; polar-error analysis | No | Spectral smoothness; bounded Frobenius variance | Average nuclear gradient; explicit rank factors |
| Choudhury et al. (2026) | Single matrix | Inexact polar, including NS; relative alignment bound | Yes | Frobenius -smoothness; gradient-dependent -moment noise, | Best-iterate Frobenius gradient; -dependent constants b |
| Li and Tsuchiya (2026) | Single-matrix online learner | Finite Taylor NS; fixed normalization | No | Nonsmooth objectives allowed; moment bounds and almost-sure operator-norm gradient bound | Online-to-nonconvex stationarity; rank-dependent bounds |