The Row Normalization Puzzle in Muon
Organizations: Department of Industrial Engineering and Operations Research Columbia University
Abstract
This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performance (Li et al.). Despite its growing adoption and promising performance in large language model (LLM) pretraining, NorMuon's worst-case guarantees remain poorly understood. One fundamental question is: Does row normalization yield provable convergence gains, potentially through its interaction with approximate polar computation and exponential moving-average momentum? Our results show that row normalization introduces a dimension-dependent factor in the worst-case iteration complexity under the operator-norm geometry, which persists even with exact polar computation and any fixed momentum parameters. Indeed, we establish an algorithm-dependent lower bound and a matching upper bound in deterministic settings, and extend our upper bound analysis to stochastic settings. Both upper-bound analyses allow approximate polar computation. Experiments show that NorMuon is slower than Muon on synthetic problems inspired by our worst-case construction, yet outperforms Muon in LLM pretraining. These findings sharpen the puzzle of why row normalization helps in practice and complement the recent findings of Dewulf et al.
Figures & tables
| Muon | |||||
|---|---|---|---|---|---|
| NorMuon (Algorithm 1 ) | |||||
| NorMuon with Eq. ( 4.1 ) |
| AdamW | Muon | NorMuon (column) | NorMuon (row) | |
|---|---|---|---|---|
| Test loss | ||||
| Test accuracy |
| MMLU | HSwag | PIQA | Wino. | ARC-C | ARC-E | BoolQ | CSQA | SIQA | OBQA | Avg | Loss | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| B | AdamW | ||||||||||||
| Muon | |||||||||||||
| NorMuon | |||||||||||||
| B | AdamW | ||||||||||||
| Muon | |||||||||||||
| NorMuon |
| NorMuon parameter groups | M | B | Other optimizers | M | B |
|---|---|---|---|---|---|
| Muon | Muon | ||||
| NorMuon | NorMuon | ||||
| QKVO only | Aurora | ||||
| MLP input only | Muon+ | ||||
| MLP output only | MuonEq-R | ||||
| NorMuon (row-wise MLP) |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Muon | NorMuon (Algorithm 1 ) | NorMuon with Eq. ( 4.1 ) | |
|---|---|---|---|
| Muon | NorMuon (Algorithm 1 ) | NorMuon with Eq. ( 4.1 ) | |
|---|---|---|---|
| Muon | |||||
|---|---|---|---|---|---|
| NorMuon (Algorithm 1 ) | |||||
| NorMuon with Eq. ( 4.1 ) |
| Model | Depth | Hidden dim. | MLP dim. | Vocab. size | Seq. length | Matrix params | Batch size (tokens) | Training tokens |
|---|---|---|---|---|---|---|---|---|
| M | M | B | ||||||
| B | M | B | ||||||
| B | B | B |
| Benchmark | Muon | NorMuon |
|---|---|---|
| MMLU | 33.3 ±0.3 | 33.3 ±0.06 |
| HellaSwag | 60.3 ±0.2 | 60.9 ±0.4 |
| PIQA | 75.9 ±0.3 | 76.1 ±0.4 |
| WinoGrande | 55.5 ±0.4 | 58.4 ±1.5 |
| ARC-C | 43.8 ±0.8 | 45.4 ±0.3 |
| ARC-E | 73.6 ±0.3 | 75.0 ±0.8 |
| Step size | ||||
|---|---|---|---|---|
| Muon | ||||
| NorMuon |
| PE steps | |||
|---|---|---|---|
| Muon | |||
| NorMuon |
| NorMuon | Muon | ||||
|---|---|---|---|---|---|
| Test loss |
| Weight decay | |||
|---|---|---|---|
| Muon | |||
| NorMuon |