Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
Authors: Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu
Organizations: Gaoling School of Artificial Intelligence Renmin University of China · School of Statistics and Management Shanghai University of Finance and Economics
Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish O(T−1/2) convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.
Figures & tables
Figure 1: GPT-2 Small on OpenWebText and LLaMA Models on C4. Training and validation loss for GPT-2 Small and LLaMA-{130M, 350M, 1.3B} with AdamW, Muon, MANO, NormPre-G and NormPre-L. Dark and light curves denote 50-step moving averages and raw training trajectories.
Figure 2: Geometric Workflow of the NormPre Optimizer (Algorithm 1 ).
Figure 3: Spectral Interaction Preconditioning.
Setting
Validation Loss ↓
Model
Dataset
AdamW
Muon
MANO
NormPre-G
NormPre-L
GPT-2 Small
OpenWebText
3.1444 ± 0.0054
3.1064 ± 0.0054
3.1156 ± 0.0057
3.0667 ± 0.0049
3.0826 ± 0.0034
LLaMA-130M
C4
3.1363
3.1019
3.1102
3.0737
3.0930
LLaMA-350M
C4
3.0378
2.9999
2.9946
2.9678
2.9758
LLaMA-1.3B
C4
2.9385
2.9037
2.8963
2.8571
2.8662
Qwen3-0.6B
Pile
2.8956
2.8335
2.8382
2.7967
2.8066
Table 1: Validation loss across model scales, architectures and pretraining datasets. GPT-2 Small results show mean ± standard deviation over three matched seeds, and larger models report single-run results under matched training budgets. Shaded columns indicate our variants. The best baseline and our outperforming variants are shown in underline and bold , respectively.
Model
Dataset
Optimizer
Optimizer Latency ↓
E2E Step Time ↓
Throughput ↑
Peak Memory ↓
(ms/step)
(ms/step)
(k tokens/s)
(GiB)
GPT-2 Small
OpenWebText
AdamW
2.7
3364.8
155.82
14.93
Muon
112.8
3479.8
150.67
14.61
MANO
8.8
3365.6
155.78
14.61
NormPre-G
119.3 ↑ 5.76%
3495.0 ↑ 0.44%
150.01 ↓ 0.44%
14.62
NormPre-L
69.8 ↓ 38.12%
3439.4 ↓ 1.16%
152.44 ↑ 1.17%
14.61
Table 2: Training Efficiency Comparisons. Measurements use the same training configurations as the main experiments and are averaged over 100 optimizer steps after 20 warmup steps.
Figure 4: (a) Eigenspectra before and after marginal normalization; (b) Spectral transformation percentages under global and localized preconditioning; (c) Performance-efficiency trade-off.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Model Size
Dataset
Layers
Hidden
FFN
Heads
Seq. Len.
Eff. Batch
Base Peak LR
Warmup
Steps
GPT-2 Small
124M
OpenWebText
12
768
3,072
12
1,024
512
6×10−4
2,000
10,000
LLaMA-130M
130M
C4
12
768
2,048
12
1,024
512
6×10−4
1,000
10,000
LLaMA-350M
350M
C4
24
1,024
2,736
16
1,024
512
3×10−4
1,000
10,000
LLaMA-1.3B
1.3B
C4
24
2,048
5,461
32
1,024
512
3×10−4
1,000
10,000
Qwen3-0.6B
0.6B
Pile
28
1,024
3,072
16
1,024
512
3×10−4
1,000
10,000
Qwen3-1.7B
1.7B
Pile
28
2,048
6,144
16
1,024
512
3×10−4
1,000
10,000
Appendix
Table 3: Model and Training Configurations.
Hyperparameter
AdamW
Muon
MANO
NormPre-G
NormPre-L
β1
0.9
–
–
–
–
β2
0.95
–
–
–
–
Momentum μ
–
0.95
0.95
0.95
0.95
Newton-Schulz steps
–
5
–
5
–
Weight decay
0.1
0.1
0.1
0.1
0.1
Interaction rank r
–
–
–
–
32
Appendix
Table 4: Default Optimizer Configurations. A dash indicates that the corresponding hyperparameter is not applicable.
Figure 5: Qwen3 Models on Pile. Training and validation loss for Qwen3-{0.6B, 1.7B} with AdamW, Muon, MANO, NormPre-G and NormPre-L. For training loss, dark curves denote 50-step moving averages and light curves the raw trajectories.
Model
Dataset
Optimizer
Optimizer Latency ↓
E2E Step Time ↓
Throughput ↑
Peak Memory ↓
(ms/step)
(ms/step)
(k tokens/s)
(GiB)
Qwen3-1.7B
Pile
AdamW
18.3
12102.7
43.32
53.17
Muon
940.0
13050.4
40.17
53.17
MANO
118.7
12217.4
42.91
53.17
NormPre-G
1003.6 ↑ 6.77%
13119.7 ↑ 0.53%
39.96 ↓ 0.52%
53.17
NormPre-L
402.7 ↓ 57.16%
12501.4 ↓ 4.21%
41.94 ↑ 4.41%
53.17
Appendix
Table 5: Training Efficiency on Qwen3-1.7B. Measurements use the same training configuration as the main Qwen3-1.7B experiment and are averaged over 100 optimizer steps after 20 warmup steps. We report global throughput across 4 GPUs and per-GPU peak allocated memory.
Figure 6: Spectral Dynamics. (a) Eigenspectra before and after marginal normalization in the column-active orientation. (b) Spectral transformation percentages across modes under global and localized preconditioning (column-active orientation).
Figure 7: Training Dynamics. (a) Pre-clipping global gradient norm over all trainable parameters. (b) Descent alignment between the gradient and optimizer update direction over matrix-optimized parameters. Dark curves show 50-step moving averages and light curves show the corresponding raw trajectories.
Variant
Validation Loss
NormPre-G
NormPre-L
No Tangent
3.0945
3.1139
Strict Tangent
3.1869
3.1196
Relaxed Tangent
3.0711
3.0864
Appendix
Table 6: Ablations on Marginally Normalized Update. Results are reported on GPT-2 Small with OpenWebText using seed 1337 under the same 10 k-step training budget.
Interaction Source
NormPre-G
NormPre-L
Training Loss
Validation Loss
Training Loss
Validation Loss
Update before normalization ( X )
3.0844
3.0900
3.1375
3.1305
Update after normalization ( Ψ )
3.0655
3.0711
3.0810
3.0864
Appendix
Table 7: Ablations on Interaction Geometry Source. Results on GPT-2 Small with OpenWebText using seed 1337 and the same 10k-step training budget.
Rank r
Training Loss
Validation Loss
8
3.0911
3.0962
16
3.0857
3.0908
32 (default)
3.0810
3.0864
64
3.0768
3.0827
Appendix
Table 8: Spectral Budget in NormPre-L. Results are reported on GPT-2 Small with OpenWebText using seed 1337 under the same 10 k-step training budget.
Method
Computational Complexity
Muon
O(mn+qmns)
NormPre-G
O(mn+qmns)
NormPre-L (Exact)
O(m2n+m3)
NormPre-L (Sketch)
O((p+1)mnℓ+(m+n)ℓ2+ℓ3)
Appendix
Table 9: Computational Complexity of Muon and NormPre Variants.
Symbol
Dimension
Description
Ψ
m×n
Marginally normalized update
Ω
n×ℓ
Gaussian random matrix
Y
m×ℓ
Gaussian sketch
Q
m×ℓ
Candidate subspace
ΓQ
ℓ×ℓ
Rayleigh–Ritz matrix
Z
ℓ×r
Top- r eigenvectors of ΓQ
Appendix
Table 10: Dimensions of the Main Quantities in Algorithm 2
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
Chang-Wei Shi, Xu Wang, Wu-Jun Li
National Key Laboratory for Novel Software Technology, School of Computer Science, Nanjing University, P. R. China
Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon mitigates gradient anisotropy, it does not explicitly account for the curvature geometry of the loss landscape and may therefore remain sensitive to curvature anisotropy. We bridge this gap by proposing MALT (Muon Augmented by Lightweight Two-sided Preconditioning), which uses lightweight diagonal preconditioners to reduce the sensitivity of Muon to curvature anisotropy. Specifically, MALT uses two-sided diagonal preconditioners with low memory and computational overhead to approximately capture the curvature geometry of the loss landscape. It orthogonalizes the preconditioned momentum using Newton-Schulz iterations and maps the result back to define the update direction, while norm grafting controls the update magnitude. To improve the robustness of MALT to stochastic gradient noise, we further propose MALTER (MALT with Adaptive stEpsize Rescaling). Convergence guarantees are provided for MALT in the stochastic non-convex setting. Experiments on GPT-2 Small, Medium, and Large pretraining show that the proposed methods outperform Muon while maintaining nearly the same memory footprint and wall-clock time.
Tongle Wu, Huanyu Dong, Ying Sun +1
School of Electrical Engineering and Computer Science, The Pennsylvania State University, University Park, PA, USA · Department of Computer Science, City University of Hong Kong, Hong Kong SAR, China
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon without decoupled weight decay, the spectral norm of weight matrices drifts upward over training. Through a decomposition of the spectral norm into a row-magnitude factor and a row-coherence factor, we identify the former as the empirical driver of this drift under Muon, while the latter remains well-behaved along the trajectory. Motivated by this diagnosis, we introduce Muown, a drop-in replacement for Muon that treats the row-magnitude vector as an explicit optimizer variable, updating it under the ℓ∞ geometry induced by the decomposition, while applying Muon unchanged to the remaining direction component. We prove that Muown attains the optimal non-convex rates in both deterministic and stochastic regimes under a dual norm aligned with the underlying geometries and with a stochastic noise coefficient that empirically remains below that of Muon throughout training. Across GPT-style pre-training on FineWeb-Edu with model sizes from 124M up to 2.7B parameters, Muown improves perplexity over Muon, SOAP, AdamW, and Lion. It also widens the plateau of near-optimal learning rates across model scales, reduces sensitivity to weight decay, and avoids the spectral norm drift at negligible step-time overhead when appropriately sharded.
Kai Lion, Florian Hübler, Bingcong Li +2
Department of Computer Science, ETH Zurich, Switzerland · Department of Mathematics, School of Computation, Information and Technology; Technical University of Munich, Germany · ELLIS Institute Tübingen, MPI-IS, Tübingen AI Center, Germany