Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.
Figures & tables
Figure 1
Figure 1: Top: Comparison of Muon, MuonEq and TensorChain; Bottom: Comparison of Aurora and Aurora-LN, both setups are training on Qwen3-0.6B with a 1M-token batch at 1× Chinchilla under cosine LR. For other setups see section F.6 .
Schedule
Model
Batch
Tokens
Muon
MuonEq
TensorChain ∗
Muon 2
Aurora
Aurora-LN ∗
WSD
0.6B
1M
8.8B
2.9201
2.9151
2.9000
2.9104
2.9184
2.9116
WSD
1.7B
4M
112B
2.5310
2.5279
2.5236
2.5264
2.5301
2.5242
Cosine
0.6B
1M
8.8B
2.8639
2.8540
2.8379
2.8497
2.8624
2.8474
Cosine
0.6B
4M
35B
2.7527
2.7434
2.7344
2.7443
2.7487
2.7391
Cosine
1.7B
1M
28.1B
2.5974
2.5924
2.5857
2.5885
2.5964
2.5882
Cosine
1.7B
4M
112B
2.4823
2.4784
2.4747
2.4744
2.4810
2.4734
Table 1: Final validation loss under WSD and cosine schedules. Bold marks the best method in each row. The two methods on the right (Aurora, Aurora-LN) use Newton–Schulz twice, whereas other methods apply Newton–Schulz once. Muon 2 has an additional momentum state. (*) were introduced in Section 4 .
Model
Batch
Tokens
Aurora-LN
Muon 2
MuonEq
Aurora
Muon
0.6B
1M
8.8B
8.2
9.0
10.3
12.4
12.8
0.6B
4M
35B
5.4
7.6
7.4
8.9
10.0
1.7B
1M
28.1B
3.9
4.0
6.2
7.5
7.8
1.7B
4M
112B
-1.9
-0.5
6.7
7.4
7.8
Table 2: Token-efficiency gain of TensorChain: percentage (%) reduction in tokens needed to reach a common validation-loss target, estimated by linear interpolation between adjacent checkpoints. Positive values favor TensorChain.
Schedule
Consecutive
Stride
Greedy
All
WSD
2.9070
2.8990
2.8994
2.9000
Cosine
2.8474
2.8357
2.8370
2.8379
Table 3: TensorChain grouping ablation on Qwen3-0.6B, 1M-token batch, 1× Chinchilla. Each entry reports the best validation loss from the corresponding full learning-rate sweep.
Schedule
Method
k=1
k=2
k=3
k=4
k=5
WSD
TensorChain
2.8992
2.9003
2.9005
2.9002
2.9000
Sink–NS–Sink
2.9210
2.9218
2.9214
2.9216
2.9213
Cosine
TensorChain
2.8368
2.8381
2.8386
2.8382
2.8379
Sink–NS–Sink
2.8624
2.8630
2.8626
2.8624
2.8627
Table 4: Ablation of cross-layer normalization and iteration count k on Qwen3-0.6B (1M-token batch, 1× Chinchilla). Each entry reports final validation loss.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Replacing Newton–Schulz with cheaper axis-wise normalizations on Qwen3-0.6B after 8.8B training tokens. (a) NS5 is retained on Q,K,V,O while the FFN uses the indicated replacement. (b) A single row–column equilibration sweep is applied to the indicated parameter classes, with NS5 retained elsewhere.
Variant
Val. loss
Δ
Muon (NS5 on all matrices)
2.5905
–
NS5 on QKVO, replacement on FFN
Row–column, k=5
2.6022
+0.012
Row–column, k=1
2.6059
+0.015
Long-axis ℓ2
2.6134
+0.023
Column ℓ2
2.6136
+0.023
Appendix
Table 5: Selected Newton–Schulz replacement results on Qwen3-0.6B after 8.8B training tokens. Each entry reports the best result over the matrix learning-rate sweep. Δ denotes the difference from Muon.
Method
1×10−3
2×10−3
3×10−3
5×10−3
8×10−3
TensorChain
2.9089
2.9002
2.9000
2.9040
2.9131
Muon 2
2.9171
2.9106
2.9104
2.9157
2.9260
Aurora-LN
2.9204
2.9128
2.9116
2.9130
2.9225
AdaMuon
2.9290
2.9156
2.9142
2.9188
2.9294
MuonEq
2.9235
2.9163
2.9151
2.9186
2.9278
MuonEq-diag
2.9258
2.9184
2.9167
2.9211
2.9305
Appendix
Table 6: Full Qwen3-0.6B, 1M-token-batch, 1× Chinchilla baseline sweep under WSD. Entries are final validation losses; methods are ordered by their best validation loss, and bold marks the best learning rate for each method.
Method
5×10−3
8×10−3
12×10−3
16×10−3
20×10−3
TensorChain
2.8501
2.8409
2.8379
2.8383
2.8399
AdaMuon
2.8674
2.8609
2.8568
2.8519
2.8458
Aurora-LN
2.8614
2.8520
2.8474
2.8476
2.8482
Muon 2
2.8626
2.8543
2.8497
2.8498
2.8515
MuonEq
2.8675
2.8578
2.8540
2.8548
2.8563
MuonEq-diag
2.8708
2.8624
2.8572
2.8562
2.8567
Appendix
Table 7: Full Qwen3-0.6B, 1M-token-batch, 1× Chinchilla baseline sweep under cosine decay. Entries are final validation losses; methods are ordered by their best validation loss, and bold marks the best learning rate for each method.
Figure 3: Full Qwen3-0.6B baseline comparison.
Matrix LR
1×10−3
2×10−3
3×10−3
5×10−3
8×10−3
Final val. loss
2.9299
2.9224
2.9201
2.9251
2.9354
Appendix
Table 8: Muon learning-rate sweep for Qwen3-0.6B, 1M-token batch, 1× Chinchilla under WSD.
Matrix LR
0.5×10−3
1×10−3
2×10−3
4×10−3
8×10−3
Final val. loss
2.6878
2.6672
2.6523
2.6493
2.6589
Appendix
Table 9: Muon learning-rate sweep for Qwen3-1.7B, 4M-token batch, 1× Chinchilla under WSD.
Matrix LR
1×10−3
2×10−3
3×10−3
5×10−3
8×10−3
12×10−3
16×10−3
20×10−3
Final val. loss
2.9201
2.8993
2.8857
2.8756
2.8680
2.8639
2.8645
2.8659
Appendix
Table 10: Muon learning-rate sweep for Qwen3-0.6B, 1M-token batch, 1× Chinchilla under cosine decay.
Matrix LR
8×10−3
12×10−3
18×10−3
27×10−3
40×10−3
60×10−3
Final val. loss
2.9650
2.9568
2.9516
2.9462
2.9459
2.9468
Appendix
Table 11: Muon learning-rate sweep for Qwen3-0.6B, 4M-token batch, 1× Chinchilla under cosine decay.
Matrix LR
4×10−3
6×10−3
9×10−3
13.5×10−3
20×10−3
Final val. loss
2.6019
2.5974
2.5987
2.6042
2.6139
Appendix
Table 12: Muon learning-rate sweep for Qwen3-1.7B, 1M-token batch, 1× Chinchilla under cosine decay.
Matrix LR
5×10−3
8×10−3
10×10−3
15×10−3
20×10−3
Final val. loss
2.6228
2.6120
2.6091
2.6047
2.6049
Appendix
Table 13: Muon learning-rate sweep for Qwen3-1.7B, 4M-token batch, 1× Chinchilla under cosine decay.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.9446
2.9437
2.9397
2.9350
2.9362
2.9249
1.7B, 4M, 4×
2.5254
2.5244
2.5226
2.5208
2.5188
2.5181
Appendix
Table 14: Final training loss for the WSD trajectory experiments. Bold marks the best method in each row.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.9201
2.9184
2.9151
2.9104
2.9116
2.8990
1.7B, 4M, 4×
2.5310
2.5301
2.5279
2.5264
2.5242
2.5236
Appendix
Table 15: Final validation loss for the WSD trajectory experiments. Bold marks the best method in each row.
Figure 4: WSD trajectories for Qwen3-0.6B, 1M-token batch, 1× Chinchilla.
Figure 5: WSD trajectories for Qwen3-1.7B, 4M-token batch, 4× Chinchilla.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.8899
2.8889
2.8806
2.8771
2.8750
2.8636
0.6B, 4M, 4×
2.7934
2.7897
2.7845
2.7851
2.7804
2.7748
1.7B, 1M, 1×
2.6036
2.6017
2.5978
2.5941
2.5942
2.5922
1.7B, 4M, 4×
2.3859
2.3850
2.3814
2.3780
2.3751
2.3778
Appendix
Table 16: Final training loss for the cosine trajectory experiments. Bold marks the best method in each row.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.8639
2.8624
2.8540
2.8497
2.8474
2.8379
0.6B, 4M, 4×
2.7527
2.7487
2.7434
2.7443
2.7391
2.7344
1.7B, 1M, 1×
2.5974
2.5964
2.5924
2.5885
2.5882
2.5857
1.7B, 4M, 4×
2.4823
2.4810
2.4784
2.4744
2.4734
2.4747
Appendix
Table 17: Final validation loss for the cosine trajectory experiments. Bold marks the best method in each row.
Figure 6: Cosine trajectories for Qwen3-0.6B, 1M-token batch, 1× Chinchilla.
Figure 7: Cosine trajectories for Qwen3-0.6B, 4M-token batch, 4× Chinchilla.
Figure 8: Cosine trajectories for Qwen3-1.7B, 1M-token batch, 1× Chinchilla.
Figure 9: Cosine trajectories for Qwen3-1.7B, 4M-token batch, 4× Chinchilla.
Table 18: TensorChain grouping ablation on Qwen3-0.6B, 1M-token batch, 1× Chinchilla. Each entry reports the best point from the corresponding full sweep.
Setting
Aurora-LN
Muon 2
MuonEq
Aurora
Muon
0.6B, 1M, 1×
8.2
9.0
10.3
12.4
12.8
0.6B, 4M, 4×
5.4
7.6
7.4
8.9
10.0
1.7B, 1M, 1×
3.9
4.0
6.2
7.5
7.8
1.7B, 4M, 4×
-2.0
-0.5
6.7
7.4
7.8
Appendix
Table 19: Cosine-schedule token-efficiency gain of TensorChain. Each entry is the percentage reduction in training tokens required to reach the same validation-loss target. Positive values favor TensorChain.
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.
Chang-Wei Shi, Xu Wang, Wu-Jun Li
National Key Laboratory for Novel Software Technology, School of Computer Science, Nanjing University, P. R. China
Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.
Yuanshi Liu, Boyuan Jiang, Liang Hou +4
School of Intelligence Science and Technology, Peking University · Kling Team
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iteration as its core operation. Standard Muon applies a uniform NS schedule to all parameter matrices, overlooking possible differences in orthogonalization difficulty and its impact on performance. Through a systematic empirical study, we show that this per-matrix heterogeneity is pervasive and strongly associated with matrix geometry, which evolves dynamically across operator types, training stages, and network depths. Therefore, uniform NS schedules can lead to uneven orthogonalization quality across the model. Motivated by these findings, we propose Operator-level Adaptive Muon Orthogonalization (AMO), an observe-then-commit method that measures weight geometry by operator type early in training and then uses these signals to allocate the NS budget for the remainder of training. AMO delivers consistent improvements over uniform-schedule Muon across standard, prolonged, and continual pre-training, surpassing the strongest baseline by +0.76 on Llama3.1-1.4B and +0.51 on Qwen3-1.7B in average downstream performance of 12 evaluation tasks, with gains persisting at Llama3.1-4B scale.
Xinlin Zhuang, Panyi Ouyang, Yichen Li +7
The Chinese University of Hong Kong · Shopee · MBZUAI +3