Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
Organizations: Gaoling School of Artificial Intelligence Renmin University of China · School of Statistics and Management Shanghai University of Finance and Economics
Abstract
Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.
Figures & tables
| Setting | Validation Loss | |||||
|---|---|---|---|---|---|---|
| Model | Dataset | AdamW | Muon | MANO | NormPre-G | NormPre-L |
| GPT-2 Small | OpenWebText | 3.1444 0.0054 | 3.1064 0.0054 | 3.1156 0.0057 | 3.0667 0.0049 | 3.0826 0.0034 |
| LLaMA-130M | C4 | 3.1363 | 3.1019 | 3.1102 | 3.0737 | 3.0930 |
| LLaMA-350M | C4 | 3.0378 | 2.9999 | 2.9946 | 2.9678 | 2.9758 |
| LLaMA-1.3B | C4 | 2.9385 | 2.9037 | 2.8963 | 2.8571 | 2.8662 |
| Qwen3-0.6B | Pile | 2.8956 | 2.8335 | 2.8382 | 2.7967 | 2.8066 |
| Model | Dataset | Optimizer | Optimizer Latency | E2E Step Time | Throughput | Peak Memory |
|---|---|---|---|---|---|---|
| (ms/step) | (ms/step) | (k tokens/s) | (GiB) | |||
| GPT-2 Small | OpenWebText | AdamW | 2.7 | 3364.8 | 155.82 | 14.93 |
| Muon | 112.8 | 3479.8 | 150.67 | 14.61 | ||
| MANO | 8.8 | 3365.6 | 155.78 | 14.61 | ||
| NormPre-G | 119.3 5.76% | 3495.0 0.44% | 150.01 0.44% | 14.62 | ||
| NormPre-L | 69.8 38.12% | 3439.4 1.16% | 152.44 1.17% | 14.61 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Model Size | Dataset | Layers | Hidden | FFN | Heads | Seq. Len. | Eff. Batch | Base Peak LR | Warmup | Steps |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-2 Small | 124M | OpenWebText | 12 | 768 | 3,072 | 12 | 1,024 | 512 | 2,000 | 10,000 | |
| LLaMA-130M | 130M | C4 | 12 | 768 | 2,048 | 12 | 1,024 | 512 | 1,000 | 10,000 | |
| LLaMA-350M | 350M | C4 | 24 | 1,024 | 2,736 | 16 | 1,024 | 512 | 1,000 | 10,000 | |
| LLaMA-1.3B | 1.3B | C4 | 24 | 2,048 | 5,461 | 32 | 1,024 | 512 | 1,000 | 10,000 | |
| Qwen3-0.6B | 0.6B | Pile | 28 | 1,024 | 3,072 | 16 | 1,024 | 512 | 1,000 | 10,000 | |
| Qwen3-1.7B | 1.7B | Pile | 28 | 2,048 | 6,144 | 16 | 1,024 | 512 | 1,000 | 10,000 |
| Hyperparameter | AdamW | Muon | MANO | NormPre-G | NormPre-L |
|---|---|---|---|---|---|
| 0.9 | – | – | – | – | |
| 0.95 | – | – | – | – | |
| Momentum | – | 0.95 | 0.95 | 0.95 | 0.95 |
| Newton-Schulz steps | – | 5 | – | 5 | – |
| Weight decay | 0.1 | 0.1 | 0.1 | 0.1 | 0.1 |
| Interaction rank | – | – | – | – | 32 |
| Model | Dataset | Optimizer | Optimizer Latency | E2E Step Time | Throughput | Peak Memory |
|---|---|---|---|---|---|---|
| (ms/step) | (ms/step) | (k tokens/s) | (GiB) | |||
| Qwen3-1.7B | Pile | AdamW | 18.3 | 12102.7 | 43.32 | 53.17 |
| Muon | 940.0 | 13050.4 | 40.17 | 53.17 | ||
| MANO | 118.7 | 12217.4 | 42.91 | 53.17 | ||
| NormPre-G | 1003.6 6.77% | 13119.7 0.53% | 39.96 0.52% | 53.17 | ||
| NormPre-L | 402.7 57.16% | 12501.4 4.21% | 41.94 4.41% | 53.17 |
| Variant | Validation Loss | |
|---|---|---|
| NormPre-G | NormPre-L | |
| No Tangent | 3.0945 | 3.1139 |
| Strict Tangent | 3.1869 | 3.1196 |
| Relaxed Tangent | 3.0711 | 3.0864 |
| Interaction Source | NormPre-G | NormPre-L | ||
|---|---|---|---|---|
| Training Loss | Validation Loss | Training Loss | Validation Loss | |
| Update before normalization ( ) | 3.0844 | 3.0900 | 3.1375 | 3.1305 |
| Update after normalization ( ) | 3.0655 | 3.0711 | 3.0810 | 3.0864 |
| Rank | Training Loss | Validation Loss |
|---|---|---|
| 8 | 3.0911 | 3.0962 |
| 16 | 3.0857 | 3.0908 |
| 32 (default) | 3.0810 | 3.0864 |
| 64 | 3.0768 | 3.0827 |
| Method | Computational Complexity |
|---|---|
| Muon | |
| NormPre-G | |
| NormPre-L (Exact) | |
| NormPre-L (Sketch) |
| Symbol | Dimension | Description |
|---|---|---|
| Marginally normalized update | ||
| Gaussian random matrix | ||
| Gaussian sketch | ||
| Candidate subspace | ||
| Rayleigh–Ritz matrix | ||
| Top- eigenvectors of |