Muon Sublates the Edge of Stability in LLM Pretraining
Authors: Yanzhe Chen, Qifang Zhao, Xiaoxiao Xu, Fanghui Liu
Organizations: School of Mathematical Sciences, Shanghai Jiao Tong University, China · Alibaba Inc., Hangzhou, China · School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University, China
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary 2ρb/η, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
Figures & tables
Figure 1: Results on a controlled Transformer (a,b,c) to language-model pretraining (d,e,f). Transformer experiments (on the top) use SVD-polar updates at a fixed learning rate η=0.007 and varying batch sizes; 130M Llama-like LLM pre-training (on the bottom) uses no-momentum NS-5 Muon at a warmup-stable-decay (WSD) ( Hu et al., 2024 ) learning rate schedule (with a peak learning rate 0.02 ). Columns show the loss, conditional effective curvature and its coherence-corrected loss-neutral boundary 2ρ^k,b/ηk , and global Muon-direction cosine in a range of [−1,1] , respectively. In the 130M LLM run, conditional loss balance coexists with continued learning and a weak negative directional bias. The inset in (e) enlarges steps 10–490. The experiments on 1B LLM pre-training are deferred to Section 3.3 .
Figure 2: Loss balance and update alignment across batch sizes. Results are shown for MLP and CNN experiments with SVD polar updates at η=0.007 . (a) and (b) compare conditional curvature s^bM with 2ρ^b/η on MLP and CNN, respectively; (c) and (d) show global consecutive-update cosines on MLP and CNN, respectively. Dark and light curves in (a) and (b) denote curvature and boundary, respectively. The corresponding results on Transformer appear in Figures 1b and 1c .
Figure 3: Diagnostic events in controlled networks. These experiments use the MLP, CNN, and Transformer at η=0.007 with the listed batch sizes; they measure the first loss-rise, T1, and T2 events. Each bar reports the first qualifying step under the recorded tolerances.
Figure 4: Loss balance and update alignment across learning rates on 130M LLM pre-training. The 130M runs use no-momentum NS-5 Muon and batch size 32. Columns show peak learning rates 0.04 , 0.08 , and 0.12 , respectively. The top row compares conditional curvature with 2ρ^k,32/ηk ; the bottom row shows global direction cosines. Figures 1e and 1f shows the learning rate 0.02 baseline.
Figure 5: Probe-batch effects on T1 quantities. This experiment uses the 130M Llama-like LLM with training batch size 32; it measures reference-gradient coherence for probe batches b=1,8,16,32 at fixed checkpoints. (a)–(d) Results for peak Muon learning rates 0.02 , 0.04 , 0.08 , and 0.12 , respectively.
Figure 6: Modulewise alignment. Median cpol during the constant-learning-rate stage of the 130M run with peak rate 0.02 .
Figure 7: Loss and update directions after learning-rate changes. Panels (a,b) show validation loss and global Muon-direction cosines ck,bpol for the 22M model; panels (c,d) show the corresponding results for the 130M Llama-like model. Shaded grey regions mark learning-rate decay.
Figure 8: Loss and update alignment in 1B pretraining. (a) shows training and validation loss; (b) shows the global Muon direction cosine ckpol ; (c) shows layer-level direction cosines. Each curve in (c) represents one layer.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Architecture
Parameters
MLP
Flattened input, 3072→128→128→10 ; tanh after the two hidden layers; biases enabled
411,146
CNN
3×3 padded convolutions, 3→32→64 channels; tanh and 2×2 average pooling after each; 4,096→128→10 dense layers with hidden tanh; biases enabled
545,098
Transformer
Two pre-LayerNorm blocks, width 64, four heads, FFN width 256 with GELU; separate Q/K/V; learned token and position embeddings; final LayerNorm and untied bias-free output head
112,512
Appendix
Table 1: Small-network architectures and parameter counts.
Figure 9: Batch-dependent loss and direction diagnostics. These experiments use the MLP, CNN, and Transformer at η=0.007 ; they measure full-objective loss, T1, and T2 across four batches. Rows correspond to MLP, CNN, and Transformer, while columns show loss, T1, and T2. T1 pairs dark curvature curves with light boundaries; curves connect saved observations.
Figure 10: Layerwise and global directional alignment. These experiments use full-batch MLP, CNN, and Transformer models at η=0.007 ; they measure block-level and global update-direction cosines. Columns show MLP, CNN, and Transformer. The top row contains layer or block trajectories, and the bottom row contains global cosines computed from aggregated inner products and norms.
Figure 11: Frozen-state and temporal direction alignment in a Tiny Transformer trained with no-momentum NS-5 Muon at batch size eight. Orange markers estimate cksame from all 120 unordered pairs of 16 independent probe batches while holding the model state fixed; bars show 95% delete-one-batch jackknife intervals. Blue curves show the realized consecutive-step statistic cktime . Both statistics use only the directions on the Muon parameter subspace. Dashed and dotted vertical lines mark zero and the warmup–stable–decay boundaries, respectively.
η
b
s^M
2ρ^/η
d
95% interval for d
0.02
16
17.182
34.453
−0.01967
[−0.01999,−0.01936]
0.04
16
17.169
17.227
−0.00026
[−0.00138,−0.00086]
0.08
16
16.964
8.613
−0.15219
[−0.14650,−0.15789]
0.04
8
13.830
14.575
−0.00340
[−0.00512,−0.00167]
0.04
32
20.321
19.406
−0.00417
[−0.00349,−0.00485]
Appendix
Table 2: Conditional loss response at checkpoint 4000. Learning rates are Muon rates; the probe batch equals the branch training batch. All increments evaluate a Muon-only virtual update on the same finite reference objective, with auxiliary parameters fixed.
Figure 12: Responses to checkpoint batch-size changes. This one-seed experiment uses a 22M model with Muon peak learning rate 0.04 and branch training batches 8 , 16 , and 32 . Panels show fixed-reference loss, conditional T1 curvature and boundary with probe batch size 16, and matrix-averaged consecutive-direction cosines. The displayed range is steps 3500–8000, with branching and schedule-change markers shown.
Figure 13: 22M loss balance after learning-rate changes. Muon rates are 0.02 , 0.04 , and 0.08 . Curves compare conditional curvature with 2ρ^k/ηk at probe batch size 16. Training batch size stays at 16.
Linear warmup fraction 0.05; constant stage; decay fraction 0.2; final LR fraction 0.1
Training / data seed
1337 / 2026
Diagnostic / evaluation seed
314159 / 271828
Appendix
Table 3: Recorded 130M pre-training configuration.
Figure 14: Training and validation loss across learning rates. This experiment uses the 130M Llama-like LLM with no-momentum NS-5 Muon at four peak learning rates; it measures training and validation loss. (a)–(d) Curves for peak rates 0.02 , 0.04 , 0.08 , and 0.12 .
Figure 15: Reference-loss checks at peak rate 0.02 . (a) Conditional Muon-only increments with pointwise 95% intervals. (b) Realized Muon-only and complete-update increments. (c) Complete-update curvature skfull and boundary 2ρkfull/ηk . (b) and (c) include dense early probes; the inset covers steps 10–490.
Figure 16: Reference-loss checks at peak rate 0.04 . (a) Conditional Muon-only increments with pointwise 95% intervals. (b) Realized Muon-only and complete-update increments. (c) Complete-update curvature skfull and boundary 2ρkfull/ηk . (b) and (c) include dense early probes; the inset covers steps 10–490.
Figure 17: Reference-loss checks at peak rate 0.08 . (a) Conditional Muon-only increments with pointwise 95% intervals. (b) Realized Muon-only and complete-update increments. (c) Complete-update curvature skfull and boundary 2ρkfull/ηk . (b) and (c) include dense early probes; the inset covers steps 10–490.
Figure 18: Reference-loss checks at peak rate 0.12 . (a) Conditional Muon-only increments with pointwise 95% intervals. (b) Realized Muon-only and complete-update increments. (c) Complete-update curvature skfull and boundary 2ρkfull/ηk . (b) and (c) include dense early probes; the inset covers steps 10–490.
Figure 19: Layer and module variation in directional alignment. This experiment uses the 130M Llama-like LLM at peak Muon learning rate 0.02 . (a) shows block trajectories; (b) shows the global cosine. Both cover the constant-learning-rate stage. Figure 6 shows module medians.
Item
Setting
Model
Llama-like; approximately 1B parameters; 16 layers
Model width / heads
2,048 / 16; head dimension 128
FFN
SwiGLU; intermediate width 6,400
Normalization / positions
RMSNorm / RoPE
Vocabulary / output head
50,304; tied input and output embeddings
Data
GPT-2-tokenized FineWeb sample-100BT
Appendix
Table 4: 1B pretraining configuration and recorded diagnostics.
School of Electrical Engineering and Computer Science, The Pennsylvania State University, University Park, PA, USA · Department of Computer Science, City University of Hong Kong, Hong Kong SAR, China