MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
Organizations: National Key Laboratory for Novel Software Technology, School of Computer Science, Nanjing University, P. R. China
Abstract
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
Figures & tables
| Parameter matrix | Shape | Row CV | Column CV |
|---|---|---|---|
| q_proj | 0.194 | 0.027 | |
| k_proj | 0.180 | 0.024 | |
| v_proj | 0.195 | 0.041 | |
| o_proj | 0.014 | 0.091 | |
| gate_proj | 0.161 | 0.017 | |
| up_proj | 0.159 | 0.016 |
| Parameter matrix | Shape | Row CV | Column CV |
|---|---|---|---|
| embed_tokens | 1.225 | 0.185 | |
| lm_head | 1.384 | 0.493 |
| Before | After | |||
| Parameter matrix | Row CV | Column CV | Row CV | Column CV |
| q_proj | 0.194 | 0.027 | 0.027 | |
| k_proj | 0.180 | 0.024 | 0.023 | |
| v_proj | 0.195 | 0.041 | 0.041 | |
| o_proj | 0.014 | 0.091 | 0.014 | |
| gate_proj | 0.161 | 0.017 | 0.018 | |
| Llama | SmolLM2 | Qwen2 | ||||
| Model scale | 60M | 130M | 350M | 135M | 360M | 0.5B |
| Token budget | 1.2B | 2.7B | 7.4B | 2.7B | 7.2B | 9.9B |
| AdamW | 37.40 | 24.46 | 17.02 | 24.88 | 18.46 | 19.65 |
| SCALE | 58.38 | 33.53 | 19.87 | – | – | – |
| Muon | 29.88 | 21.64 | 16.04 | 22.81 | 17.28 | 18.79 |
| NorMuon | 29.72 | 21.57 | 15.98 | 22.83 | 17.20 | 18.76 |
| Llama | SmolLM2 | Qwen2 | ||||
|---|---|---|---|---|---|---|
| 60M | 130M | 350M | 135M | 360M | 0.5B | |
| Muon | 346.57 | 699.15 | 1653.88 | 621.27 | 1560.48 | 2404.17 |
| NorMuon | 346.73 | 699.51 | 1654.85 | 621.86 | 1561.53 | 2405.33 |
| MeqMuon | 221.53 | 511.57 | 1403.69 | 513.13 | 1380.24 | 1884.59 |
| Llama-350M | SmolLM2-360M | |||||
|---|---|---|---|---|---|---|
| Module | Shape | Row (%) | Col. (%) | Shape | Row (%) | Col. (%) |
| q_proj | 89.87 | 10.13 | 94.02 | 5.98 | ||
| k_proj | 99.21 | 0.79 | 3.97 | 96.03 | ||
| v_proj | 99.72 | 0.28 | 1.80 | 98.20 | ||
| o_proj | 10.06 | 89.94 | 18.85 | 81.15 | ||
| gate_proj | 100.00 | 0.00 | 100.00 | 0.00 | ||
| Llama-60M | Llama-130M | |||
|---|---|---|---|---|
| Variant | PPL ( ) | Memory (MiB, ) | PPL ( ) | Memory (MiB, ) |
| Muon | 29.88 | 346.57 | 21.64 | 699.15 |
| Only 2D weights in hidden layers | 29.74 | 346.57 | 21.53 | 699.15 |
| Only remaining 2D parameters | 29.81 | 221.57 | 21.52 | 511.65 |
| Only 1D parameters | 29.81 | 346.53 | 21.69 | 699.07 |
| MeqMuon | 29.53 | 221.53 | 21.38 | 511.57 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Step 381 | Step 1144 | Step 1907 | ||||||
|---|---|---|---|---|---|---|---|---|
| Layer | Matrix | Shape | ||||||
| 0 | q_proj | 0.033 | 0.012 | 0.084 | 0.022 | 0.066 | 0.024 | |
| k_proj | 0.026 | 0.012 | 0.079 | 0.022 | 0.073 | 0.027 | ||
| v_proj | 0.078 | 0.019 | 0.084 | 0.027 | 0.082 | 0.024 | ||
| o_proj | 0.015 | 0.063 | 0.013 | 0.053 | 0.013 | 0.048 | ||
| gate_proj | 0.603 | 0.020 | 0.534 | 0.026 | 0.528 | 0.023 | ||
| Step 443 | Step 1329 | Step 2215 | ||||||
|---|---|---|---|---|---|---|---|---|
| Layer | Matrix | Shape | ||||||
| 0 | q_proj | 0.034 | 0.019 | 0.043 | 0.023 | 0.040 | 0.021 | |
| k_proj | 0.088 | 0.031 | 0.076 | 0.031 | 0.066 | 0.026 | ||
| v_proj | 0.164 | 0.081 | 0.101 | 0.041 | 0.066 | 0.025 | ||
| o_proj | 0.034 | 0.144 | 0.020 | 0.087 | 0.017 | 0.050 | ||
| gate_proj | 0.654 | 0.022 | 0.528 | 0.017 | 0.503 | 0.017 | ||
| Step 381 | Step 1144 | Step 1907 | ||||||
|---|---|---|---|---|---|---|---|---|
| Layer | Matrix | Shape | ||||||
| 0 | q_proj | 0.018 | 0.014 | 0.033 | 0.015 | 0.026 | 0.013 | |
| k_proj | 0.023 | 0.084 | 0.027 | 0.089 | 0.030 | 0.090 | ||
| v_proj | 0.020 | 0.082 | 0.020 | 0.084 | 0.020 | 0.082 | ||
| o_proj | 0.019 | 0.038 | 0.016 | 0.027 | 0.015 | 0.026 | ||
| gate_proj | 0.421 | 0.023 | 0.370 | 0.021 | 0.363 | 0.019 | ||
| Family | Size | Params (M) | Layers | Hidden | FFN | Q/KV | Head dim | Vocab. | Tied |
|---|---|---|---|---|---|---|---|---|---|
| Llama | 60M | 58.074 | 8 | 512 | 1376 | 8/8 | 64 | 32,000 | No |
| 130M | 134.106 | 12 | 768 | 2048 | 12/12 | 64 | 32,000 | No | |
| 350M | 367.969 | 24 | 1024 | 2736 | 16/16 | 64 | 32,000 | No | |
| SmolLM2 | 135M | 134.515 | 30 | 576 | 1536 | 9/3 | 64 | 49,152 | Yes |
| 360M | 361.821 | 32 | 960 | 2560 | 15/5 | 64 | 49,152 | Yes | |
| Qwen2 | 0.5B | 494.033 | 24 | 896 | 4864 | 14/2 | 64 | 151,936 | Yes |