cs.LGOct 8, 2026

From Geometry to Generalization: Why Row Normalization Can Beat Adam and Muon

Authors: Jihwan Kim, Dogyoon Song, Chulhee Yun

Organizations: Seoul National University · KAIST InnoCORE LLM · University of California, Davis · KAIST

Abstract

Different optimizers can fit the same training data while selecting classifiers with substantially different geometries, but whether this difference provably affects population performance remains unclear. We show that row-wise normalization can achieve strictly higher population accuracy than full-batch Adam, a proxy for random-reshuffling Adam, and exact-SVD Muon in high-dimensional multiclass classification. Under an isotropic Gaussian-cloud data model, this advantage arises because row normalization's class-wise Euclidean geometry asymptotically preserves the population decision-boundary directions, whereas Adam's coordinate-wise geometry and Muon's spectral geometry introduce nonvanishing distortions. Beyond isotropy, the advantage persists for full-batch training on class means with independently oriented class-mean and test-noise covariances. It holds for power-law spectra with class-mean exponent below one, even under heavily anisotropic test noise. When both covariances are diagonal and sufficiently close, the advantage over Adam can reverse, while applying the same random rotation to both restores it by changing only their alignment with Adam's coordinate axes. Synthetic and last-layer language-model experiments support the predicted advantage.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Muown: Row-Norm Control for Muon Optimization

    May 11, 2026Kai Lion, Florian Hübler, Bingcong Li +2Language Model PretrainingMuon Optimizer

  2. The Row Normalization Puzzle in Muon

    Sep 30, 2026Jiayu Zhang, Tianyi LinStochastic OptimizationOptimization Convergence Analysis

  3. On the Two Faces of Adam in Separable Linear Classification

    Sep 27, 2026Chen Fan, Csaba SzepesváriImplicit BiasClassification