Muon Sublates the Edge of Stability in LLM Pretraining
Organizations: School of Mathematical Sciences, Shanghai Jiao Tong University, China · Alibaba Inc., Hangzhou, China · School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University, China
Abstract
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary , while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
Figures & tables
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Architecture | Parameters |
|---|---|---|
| MLP | Flattened input, ; tanh after the two hidden layers; biases enabled | 411,146 |
| CNN | padded convolutions, channels; tanh and average pooling after each; dense layers with hidden tanh; biases enabled | 545,098 |
| Transformer | Two pre-LayerNorm blocks, width 64, four heads, FFN width 256 with GELU; separate Q/K/V; learned token and position embeddings; final LayerNorm and untied bias-free output head | 112,512 |
| 95% interval for | |||||
|---|---|---|---|---|---|
| 0.02 | 16 | 17.182 | 34.453 | ||
| 0.04 | 16 | 17.169 | 17.227 | ||
| 0.08 | 16 | 16.964 | 8.613 | ||
| 0.04 | 8 | 13.830 | 14.575 | ||
| 0.04 | 32 | 20.321 | 19.406 |
| Item | Setting |
|---|---|
| Muon | NS-5; momentum 0; weight decay 0; update scale 1 |
| Muon peak learning rates | 0.02, 0.04, 0.08, 0.12 |
| Auxiliary optimizer | AdamW; peak LR 0.001; betas ; ; weight decay 0.1 |
| Schedule | Linear warmup fraction 0.05; constant stage; decay fraction 0.2; final LR fraction 0.1 |
| Training / data seed | 1337 / 2026 |
| Diagnostic / evaluation seed | 314159 / 271828 |
| Item | Setting |
|---|---|
| Model | Llama-like; approximately 1B parameters; 16 layers |
| Model width / heads | 2,048 / 16; head dimension 128 |
| FFN | SwiGLU; intermediate width 6,400 |
| Normalization / positions | RMSNorm / RoPE |
| Vocabulary / output head | 50,304; tied input and output embeddings |
| Data | GPT-2-tokenized FineWeb sample-100BT |