The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
Figures & tables
Parameter matrix
Shape
Row CV
Column CV
q_proj
1024×1024
0.194
0.027
k_proj
1024×1024
0.180
0.024
v_proj
1024×1024
0.195
0.041
o_proj
1024×1024
0.014
0.091
gate_proj
2736×1024
0.161
0.017
up_proj
2736×1024
0.159
0.016
Table 1: Row and column CVs of Muon’s orthogonalized updates after 5 NS iterations.
Parameter matrix
Shape
Row CV
Column CV
embed_tokens
32000×1024
1.225
0.185
lm_head
32000×1024
1.384
0.493
Table 2: Row and column CVs of unorthogonalized momentum matrices for the remaining 2D parameters.
Before
After
Parameter matrix
Row CV
Column CV
Row CV
Column CV
q_proj
0.194
0.027
<0.001
0.027
k_proj
0.180
0.024
<0.001
0.023
v_proj
0.195
0.041
<0.001
0.041
o_proj
0.014
0.091
0.014
<0.001
gate_proj
0.161
0.017
<0.001
0.018
Table 3: Row and column CVs before and after normalization.
Llama
SmolLM2
Qwen2
Model scale
60M
130M
350M
135M
360M
0.5B
Token budget
1.2B
2.7B
7.4B
2.7B
7.2B
9.9B
AdamW
37.40
24.46
17.02
24.88
18.46
19.65
SCALE
58.38
33.53
19.87
–
–
–
Muon
29.88
21.64
16.04
22.81
17.28
18.79
NorMuon
29.72
21.57
15.98
22.83
17.20
18.76
Table 4: Validation PPL for pretraining different models ( ↓ ).
Llama
SmolLM2
Qwen2
60M
130M
350M
135M
360M
0.5B
Muon
346.57
699.15
1653.88
621.27
1560.48
2404.17
NorMuon
346.73
699.51
1654.85
621.86
1561.53
2405.33
MeqMuon
221.53
511.57
1403.69
513.13
1380.24
1884.59
Table 5: Optimizer-state memory of different optimizers (MiB, ↓ ).
Llama-350M
SmolLM2-360M
Module
Shape
Row (%)
Col. (%)
Shape
Row (%)
Col. (%)
q_proj
1024×1024
89.87
10.13
960×960
94.02
5.98
k_proj
1024×1024
99.21
0.79
320×960
3.97
96.03
v_proj
1024×1024
99.72
0.28
320×960
1.80
98.20
o_proj
1024×1024
10.06
89.94
960×960
18.85
81.15
gate_proj
2736×1024
100.00
0.00
2560×960
100.00
0.00
Table 6: Row-wise and column-wise normalization selections during pretraining.
Llama-60M
Llama-130M
Variant
PPL ( ↓ )
Memory (MiB, ↓ )
PPL ( ↓ )
Memory (MiB, ↓ )
Muon
29.88
346.57
21.64
699.15
Only 2D weights in hidden layers
29.74
346.57
21.53
699.15
Only remaining 2D parameters
29.81
221.57
21.52
511.65
Only 1D parameters
29.81
346.53
21.69
699.07
MeqMuon
29.53
221.53
21.38
511.57
Table 7: Component-wise ablation of validation PPL and optimizer-state memory on Llama models.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Step 381
Step 1144
Step 1907
Layer
Matrix
Shape
γr
γc
γr
γc
γr
γc
0
q_proj
1024×1024
0.033
0.012
0.084
0.022
0.066
0.024
k_proj
1024×1024
0.026
0.012
0.079
0.022
0.073
0.027
v_proj
1024×1024
0.078
0.019
0.084
0.027
0.082
0.024
o_proj
1024×1024
0.015
0.063
0.013
0.053
0.013
0.048
gate_proj
2736×1024
0.603
0.020
0.534
0.026
0.528
0.023
Appendix
Table 8: Row and column CVs for Llama-350M after 5 NS iterations.
Step 443
Step 1329
Step 2215
Layer
Matrix
Shape
γr
γc
γr
γc
γr
γc
0
q_proj
512×512
0.034
0.019
0.043
0.023
0.040
0.021
k_proj
512×512
0.088
0.031
0.076
0.031
0.066
0.026
v_proj
512×512
0.164
0.081
0.101
0.041
0.066
0.025
o_proj
512×512
0.034
0.144
0.020
0.087
0.017
0.050
gate_proj
1376×512
0.654
0.022
0.528
0.017
0.503
0.017
Appendix
Table 9: Row and column CVs for Llama-60M after 5 NS iterations.
Step 381
Step 1144
Step 1907
Layer
Matrix
Shape
γr
γc
γr
γc
γr
γc
0
q_proj
576×576
0.018
0.014
0.033
0.015
0.026
0.013
k_proj
192×576
0.023
0.084
0.027
0.089
0.030
0.090
v_proj
192×576
0.020
0.082
0.020
0.084
0.020
0.082
o_proj
576×576
0.019
0.038
0.016
0.027
0.015
0.026
gate_proj
1536×576
0.421
0.023
0.370
0.021
0.363
0.019
Appendix
Table 10: Row and column CVs for SmolLM2-135M after 5 NS iterations.
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon without decoupled weight decay, the spectral norm of weight matrices drifts upward over training. Through a decomposition of the spectral norm into a row-magnitude factor and a row-coherence factor, we identify the former as the empirical driver of this drift under Muon, while the latter remains well-behaved along the trajectory. Motivated by this diagnosis, we introduce Muown, a drop-in replacement for Muon that treats the row-magnitude vector as an explicit optimizer variable, updating it under the ℓ∞ geometry induced by the decomposition, while applying Muon unchanged to the remaining direction component. We prove that Muown attains the optimal non-convex rates in both deterministic and stochastic regimes under a dual norm aligned with the underlying geometries and with a stochastic noise coefficient that empirically remains below that of Muon throughout training. Across GPT-style pre-training on FineWeb-Edu with model sizes from 124M up to 2.7B parameters, Muown improves perplexity over Muon, SOAP, AdamW, and Lion. It also widens the plateau of near-optimal learning rates across model scales, reduces sensitivity to weight decay, and avoids the spectral norm drift at negligible step-time overhead when appropriately sharded.
Kai Lion, Florian Hübler, Bingcong Li +2
Department of Computer Science, ETH Zurich, Switzerland · Department of Mathematics, School of Computation, Information and Technology; Technical University of Munich, Germany · ELLIS Institute Tübingen, MPI-IS, Tübingen AI Center, Germany
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iteration as its core operation. Standard Muon applies a uniform NS schedule to all parameter matrices, overlooking possible differences in orthogonalization difficulty and its impact on performance. Through a systematic empirical study, we show that this per-matrix heterogeneity is pervasive and strongly associated with matrix geometry, which evolves dynamically across operator types, training stages, and network depths. Therefore, uniform NS schedules can lead to uneven orthogonalization quality across the model. Motivated by these findings, we propose Operator-level Adaptive Muon Orthogonalization (AMO), an observe-then-commit method that measures weight geometry by operator type early in training and then uses these signals to allocate the NS budget for the remainder of training. AMO delivers consistent improvements over uniform-schedule Muon across standard, prolonged, and continual pre-training, surpassing the strongest baseline by +0.76 on Llama3.1-1.4B and +0.51 on Qwen3-1.7B in average downstream performance of 12 evaluation tasks, with gains persisting at Llama3.1-4B scale.
Xinlin Zhuang, Panyi Ouyang, Yichen Li +7
The Chinese University of Hong Kong · Shopee · MBZUAI +3
Optimizer design plays a central role in efficient language model pretraining, directly affecting optimization dynamics, convergence speed, and compute cost under fixed training budgets. Muon has emerged as a strong optimizer by orthogonalizing momentum updates, yielding a matrix-valued analogue of sign-based normalization. However, unlike Adam-style methods, Muon does not explicitly incorporate gradient-variance information into its updates. Motivated by Adam's variance-adaptive interpretation, we propose Muon-NSR and Muon-VS, two variance-adaptive Muon variants for language model pretraining. Muon-NSR applies noise-to-signal ratio (NSR) modulation before Newton--Schulz orthogonalization, whereas Muon-VS uses variance scaling (VS) without introducing any additional hyperparameters beyond those of Muon. Both methods preserve Muon's spectral normalization structure while requiring only one additional variance buffer. Experiments on Llama-style and GPT-2 pretraining across model scales from 125M to 1.2B parameters show that our methods improve over well-tuned Muon baselines and remain competitive with representative adaptive Muon-family baselines. On Llama-1.2B, Muon-VS achieves a 1.33× step-to-target speedup over a well-tuned Muon baseline, with Muon's final validation loss as the target. These results indicate that variance-adaptive modulation is a simple and effective mechanism for improving Muon-style optimizers in language model pretraining.
Jingru Li, Yibo Fan, Huan Li
College of Artificial Intelligence, Nankai University