cs.LGOct 7, 2026

Spectrally Targeted Muon

Authors: Vishrut Goyal, Rohan Ramkumar

Organizations: Department of Computer Science, Stanford University · Department of Electrical Engineering, Stanford University.

Abstract

The Muon optimizer orthogonalizes each update matrix, setting all of its singular values to one, and has proven highly effective for training large language models. It remains unclear, however, whether this success comes from amplifying small singular directions that gradient descent neglects or from suppressing large, degenerate directions that disrupt training. We introduce Spectrally Targeted Muon, which orthogonalizes only the singular values above or below a threshold ττ, so that varying ττ interpolates between normalized SGD and Muon. It isolates the relevant singular subspaces with projections computed by Newton-Schulz iteration on a shifted Gram matrix, so no SVD is needed. We evaluate these variants on the CIFAR-10 and NanoGPT speedruns, tracking the effective rank of gradient, update, and weight matrices and a new metric, the alignment of updates with the tangent space of the weight matrix's isospectral manifold. We find three things. First, the small singular values of the momentum are not noise. Orthogonalizing everything except the few largest singular values of each matrix nearly matches Muon while touching only a small fraction of the momentum, whereas orthogonalizing only the top falls well short even though it holds almost all of it. On language models every momentum singular value is far below one, so targeted orthogonalization can only amplify, and Muon wins by making directions that are too small to train on at their raw scale trainable. Second, shrinking the largest singular values is what keeps the parameter spectrum flat; this is cheap, and it is not what drives the loss. Third, AdamW differs from Muon mainly in how slowly it builds structure, which explains its slower start and why warmup helps AdamW but only hurts Muon.

Figures & tables

Explore similar work

CardsList
  1. Reassessing Muon for Matrix Factorization

    Jul 14, 2026Ali Parviz, Gal Mishne, Alex CloningerDeep Learning OptimizationMuon Optimizer

  2. Spectral Scaling Laws of Muon

    Jun 2, 2026Gagik Magakyan, Pablo Parrilo, Asuman OzdaglarNewton-Schulz IterationMomentum Methods

  3. The Spectral Dynamics and Noise Geometry of Muon

    Jun 7, 2026Pierfrancesco Beneventano, Mahmoud Abdelmoneum, Tomaso PoggioMuon OptimizerMatrix Optimization