The Muon optimizer orthogonalizes each update matrix, setting all of its singular values to one, and has proven highly effective for training large language models. It remains unclear, however, whether this success comes from amplifying small singular directions that gradient descent neglects or from suppressing large, degenerate directions that disrupt training. We introduce Spectrally Targeted Muon, which orthogonalizes only the singular values above or below a threshold τ, so that varying τ interpolates between normalized SGD and Muon. It isolates the relevant singular subspaces with projections computed by Newton-Schulz iteration on a shifted Gram matrix, so no SVD is needed. We evaluate these variants on the CIFAR-10 and NanoGPT speedruns, tracking the effective rank of gradient, update, and weight matrices and a new metric, the alignment of updates with the tangent space of the weight matrix's isospectral manifold. We find three things. First, the small singular values of the momentum are not noise. Orthogonalizing everything except the few largest singular values of each matrix nearly matches Muon while touching only a small fraction of the momentum, whereas orthogonalizing only the top falls well short even though it holds almost all of it. On language models every momentum singular value is far below one, so targeted orthogonalization can only amplify, and Muon wins by making directions that are too small to train on at their raw scale trainable. Second, shrinking the largest singular values is what keeps the parameter spectrum flat; this is cheap, and it is not what drives the loss. Third, AdamW differs from Muon mainly in how slowly it builds structure, which explains its slower start and why warmup helps AdamW but only hurts Muon.
Figures & tables
Figure 1: Mean evaluation accuracy across 20 trials for each variant, with error bars representing ± 2 standard deviations away from the mean.
Figure 2: Tangent alignment and weight spectrum plots for CNN trained with bottom and top targeted Muon over 8 training epochs.
Figure 3: Singular values of the raw gradient and of the momentum that is orthogonalized (a; dotted lines mark the smallest and largest τ ), training loss (b), and final validation loss against the threshold (c) and against the fraction of the momentum that is orthogonalized (d; by squared Frobenius norm, matrix average). Every momentum singular value lies far below one, so the targeted band is always amplified to one; Muon and NSGD are the endpoints of the family.
τ=3×10−3
τ=10−3
τ=4×10−4
τ=2×10−4
τ=10−4
Top-targeted
3.875 ± .011
3.714 ± .002
3.573 ± .013
3.486 ± .004
3.425 ± .012
Bottom-targeted
3.309 ± .002
3.409 ± .003
3.633 ± .127 †
3.739 ± .022 †
3.772 ± .033 †
Muon
3.280 ± .008
AdamW w/ warmup
3.396 ± .004
Muon w/ warmup
3.309 ± .001
AdamW w/ warmup, wd 0
3.397 ± .004
Muon, lr 0.005
3.332
AdamW
3.468 ± .008
NSGD
3.591 ± .038 †
Top τ=3×10−3 , lr 0.49
4.800
Table 1: Final FineWeb validation loss at step 1530, mean ± half-range over seeds (three seeds for the targeted family and Muon, two for the other optimizers, one for the learning rate controls). † : at least one seed shows a loss spike; for the bottom arm these are carried by the MLP output matrices, which these thresholds never lift, when a hard batch near step 1020 passes through them un-normalized.
Figure 4: Effective rank of the gradient (a) and update (b), averaged over all matrices, and of the value (c), attention-output (d) and MLP (e) parameters, averaged over layers (seed 0). Dashed lines are the warmup variants, the dotted line in (a) marks the start of the learning rate cooldown, and the legend in (b) applies to all panels.
Figure 5: How much of the update changes the spectrum, and where. (a) Log-barrier of the isospectral tangent fraction, averaged over matrices; the grey dotted line is the value for a random rotation, the orange dotted line marks the start of the learning rate cooldown, and dashed lines are the warmup variants. (b) Share of the spectrum-changing part of the update at each singular direction of the parameter, ranked by singular value, averaged over matrices and steps 100–1500; the dotted line is uniform. (c) Final validation loss against final parameter effective rank, including the learning rate controls; the bottom-targeted runs at τ≤4×10−4 leave their MLP output matrices barely trained.
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon's advantages stem from its update rule itself and which are artifacts of the scale, architecture, and data of modern deep networks. In this work, we isolate the optimizer from these confounding factors by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled comparison against carefully tuned adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting and that several previously reported advantages are sensitive to hyperparameter choices. Our results provide a more nuanced picture of when spectrum-aware orthogonalization is beneficial and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.
Ali Parviz, Gal Mishne, Alex Cloninger
†. Halicioğlu Data Science Institute, UC San Diego. · ∗. Department of Mathematics, UC San Diego.
Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon. To keep these updates tractable, Muon performs the orthonormalization with the Newton--Schulz (NS) iteration. Since NS is only approximate, directions with small singular values fail to be orthonormalized. In Muon, NS is applied to the momentum matrix at every step, yet little is known about how the singular value spectrum of these momentum matrices behaves during training, or how that behavior changes with model size. We present the first systematic study of this question. Tracking singular value quantiles of the momentum buffer across layers in models ranging from 77M to 2.8B parameters, we observe a consistent picture: after a short burn-in, the quantiles stabilize at a value determined by the layer type and model size. These stabilization values follow remarkably clean power laws in model size, with layer-dependent exponents. Layers up to mid-late depth scale very mildly with model size M (around M−0.25), so the standard 5-step NS configuration used at academic scale will continue to orthonormalize them at much larger scales. Some of the late layers, however, scale much more aggressively (up to M−0.96) and will fall into the NS failure regime at frontier scale unless one uses more NS iterations or better-tuned coefficients. NS iterations are computationally expensive at scale; our laws give practitioners a principled, layer-aware recipe for choosing the minimum NS configuration that still orthonormalizes the directions that matter -- avoiding unnecessary computation without sacrificing update quality.
Muon replaces a matrix gradient G=UΣV⊤ by its polar factor UV⊤. This keeps the singular directions selected by the gradient, but makes the update spectrum flat. We study the optimization bias created by this operation. Under explicit alignment assumptions, we prove that the polar update is the one-step entropy-maximizing choice among bounded updates that use the gradient singular directions and do not adapt to the current weight spectrum. In an underdetermined regression model, we derive exact singular-value dynamics for continuous-time Muon and identify a measurement-dependent condition under which the normalized spectrum moves toward equal nonzero singular values. This geometry also rules out a common low-rank interpretation: at fixed Frobenius norm, Muon's distinguished state has a flat spectrum, whereas nuclear-norm minimization favors spectral concentration. Controlled matrix-sensing experiments separate the effect from simple gradient rescaling, show that norm-matched gradient descent does not reproduce Muon, and recover the predicted flattening trend across broad ablations. In small NanoGPT pretraining, Muon preserves stable rank, has a broad learning-rate plateau, and improves validation loss relative to AdamW; in a matched small-ViT control, the ranking reverses. The resulting picture is regime-dependent: Muon is not universally superior, but its flat-spectrum bias can help when many spectral directions need to remain active.