Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.
Figures & tables
Figure 1
Figure 1: Top: Comparison of Muon, MuonEq and TensorChain; Bottom: Comparison of Aurora and Aurora-LN, both setups are training on Qwen3-0.6B with a 1M-token batch at 1× Chinchilla under cosine LR. For other setups see section F.6 .
Schedule
Model
Batch
Tokens
Muon
MuonEq
TensorChain ∗
Muon 2
Aurora
Aurora-LN ∗
WSD
0.6B
1M
8.8B
2.9201
2.9151
2.9000
2.9104
2.9184
2.9116
WSD
1.7B
4M
112B
2.5310
2.5279
2.5236
2.5264
2.5301
2.5242
Cosine
0.6B
1M
8.8B
2.8639
2.8540
2.8379
2.8497
2.8624
2.8474
Cosine
0.6B
4M
35B
2.7527
2.7434
2.7344
2.7443
2.7487
2.7391
Cosine
1.7B
1M
28.1B
2.5974
2.5924
2.5857
2.5885
2.5964
2.5882
Cosine
1.7B
4M
112B
2.4823
2.4784
2.4747
2.4744
2.4810
2.4734
Table 1: Final validation loss under WSD and cosine schedules. Bold marks the best method in each row. The two methods on the right (Aurora, Aurora-LN) use Newton–Schulz twice, whereas other methods apply Newton–Schulz once. Muon 2 has an additional momentum state. (*) were introduced in Section 4 .
Model
Batch
Tokens
Aurora-LN
Muon 2
MuonEq
Aurora
Muon
0.6B
1M
8.8B
8.2
9.0
10.3
12.4
12.8
0.6B
4M
35B
5.4
7.6
7.4
8.9
10.0
1.7B
1M
28.1B
3.9
4.0
6.2
7.5
7.8
1.7B
4M
112B
-1.9
-0.5
6.7
7.4
7.8
Table 2: Token-efficiency gain of TensorChain: percentage (%) reduction in tokens needed to reach a common validation-loss target, estimated by linear interpolation between adjacent checkpoints. Positive values favor TensorChain.
Schedule
Consecutive
Stride
Greedy
All
WSD
2.9070
2.8990
2.8994
2.9000
Cosine
2.8474
2.8357
2.8370
2.8379
Table 3: TensorChain grouping ablation on Qwen3-0.6B, 1M-token batch, 1× Chinchilla. Each entry reports the best validation loss from the corresponding full learning-rate sweep.
Schedule
Method
k=1
k=2
k=3
k=4
k=5
WSD
TensorChain
2.8992
2.9003
2.9005
2.9002
2.9000
Sink–NS–Sink
2.9210
2.9218
2.9214
2.9216
2.9213
Cosine
TensorChain
2.8368
2.8381
2.8386
2.8382
2.8379
Sink–NS–Sink
2.8624
2.8630
2.8626
2.8624
2.8627
Table 4: Ablation of cross-layer normalization and iteration count k on Qwen3-0.6B (1M-token batch, 1× Chinchilla). Each entry reports final validation loss.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Replacing Newton–Schulz with cheaper axis-wise normalizations on Qwen3-0.6B after 8.8B training tokens. (a) NS5 is retained on Q,K,V,O while the FFN uses the indicated replacement. (b) A single row–column equilibration sweep is applied to the indicated parameter classes, with NS5 retained elsewhere.
Variant
Val. loss
Δ
Muon (NS5 on all matrices)
2.5905
–
NS5 on QKVO, replacement on FFN
Row–column, k=5
2.6022
+0.012
Row–column, k=1
2.6059
+0.015
Long-axis ℓ2
2.6134
+0.023
Column ℓ2
2.6136
+0.023
Appendix
Table 5: Selected Newton–Schulz replacement results on Qwen3-0.6B after 8.8B training tokens. Each entry reports the best result over the matrix learning-rate sweep. Δ denotes the difference from Muon.
Method
1×10−3
2×10−3
3×10−3
5×10−3
8×10−3
TensorChain
2.9089
2.9002
2.9000
2.9040
2.9131
Muon 2
2.9171
2.9106
2.9104
2.9157
2.9260
Aurora-LN
2.9204
2.9128
2.9116
2.9130
2.9225
AdaMuon
2.9290
2.9156
2.9142
2.9188
2.9294
MuonEq
2.9235
2.9163
2.9151
2.9186
2.9278
MuonEq-diag
2.9258
2.9184
2.9167
2.9211
2.9305
Appendix
Table 6: Full Qwen3-0.6B, 1M-token-batch, 1× Chinchilla baseline sweep under WSD. Entries are final validation losses; methods are ordered by their best validation loss, and bold marks the best learning rate for each method.
Method
5×10−3
8×10−3
12×10−3
16×10−3
20×10−3
TensorChain
2.8501
2.8409
2.8379
2.8383
2.8399
AdaMuon
2.8674
2.8609
2.8568
2.8519
2.8458
Aurora-LN
2.8614
2.8520
2.8474
2.8476
2.8482
Muon 2
2.8626
2.8543
2.8497
2.8498
2.8515
MuonEq
2.8675
2.8578
2.8540
2.8548
2.8563
MuonEq-diag
2.8708
2.8624
2.8572
2.8562
2.8567
Appendix
Table 7: Full Qwen3-0.6B, 1M-token-batch, 1× Chinchilla baseline sweep under cosine decay. Entries are final validation losses; methods are ordered by their best validation loss, and bold marks the best learning rate for each method.
Figure 3: Full Qwen3-0.6B baseline comparison.
Matrix LR
1×10−3
2×10−3
3×10−3
5×10−3
8×10−3
Final val. loss
2.9299
2.9224
2.9201
2.9251
2.9354
Appendix
Table 8: Muon learning-rate sweep for Qwen3-0.6B, 1M-token batch, 1× Chinchilla under WSD.
Matrix LR
0.5×10−3
1×10−3
2×10−3
4×10−3
8×10−3
Final val. loss
2.6878
2.6672
2.6523
2.6493
2.6589
Appendix
Table 9: Muon learning-rate sweep for Qwen3-1.7B, 4M-token batch, 1× Chinchilla under WSD.
Matrix LR
1×10−3
2×10−3
3×10−3
5×10−3
8×10−3
12×10−3
16×10−3
20×10−3
Final val. loss
2.9201
2.8993
2.8857
2.8756
2.8680
2.8639
2.8645
2.8659
Appendix
Table 10: Muon learning-rate sweep for Qwen3-0.6B, 1M-token batch, 1× Chinchilla under cosine decay.
Matrix LR
8×10−3
12×10−3
18×10−3
27×10−3
40×10−3
60×10−3
Final val. loss
2.9650
2.9568
2.9516
2.9462
2.9459
2.9468
Appendix
Table 11: Muon learning-rate sweep for Qwen3-0.6B, 4M-token batch, 1× Chinchilla under cosine decay.
Matrix LR
4×10−3
6×10−3
9×10−3
13.5×10−3
20×10−3
Final val. loss
2.6019
2.5974
2.5987
2.6042
2.6139
Appendix
Table 12: Muon learning-rate sweep for Qwen3-1.7B, 1M-token batch, 1× Chinchilla under cosine decay.
Matrix LR
5×10−3
8×10−3
10×10−3
15×10−3
20×10−3
Final val. loss
2.6228
2.6120
2.6091
2.6047
2.6049
Appendix
Table 13: Muon learning-rate sweep for Qwen3-1.7B, 4M-token batch, 1× Chinchilla under cosine decay.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.9446
2.9437
2.9397
2.9350
2.9362
2.9249
1.7B, 4M, 4×
2.5254
2.5244
2.5226
2.5208
2.5188
2.5181
Appendix
Table 14: Final training loss for the WSD trajectory experiments. Bold marks the best method in each row.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.9201
2.9184
2.9151
2.9104
2.9116
2.8990
1.7B, 4M, 4×
2.5310
2.5301
2.5279
2.5264
2.5242
2.5236
Appendix
Table 15: Final validation loss for the WSD trajectory experiments. Bold marks the best method in each row.
Figure 4: WSD trajectories for Qwen3-0.6B, 1M-token batch, 1× Chinchilla.
Figure 5: WSD trajectories for Qwen3-1.7B, 4M-token batch, 4× Chinchilla.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.8899
2.8889
2.8806
2.8771
2.8750
2.8636
0.6B, 4M, 4×
2.7934
2.7897
2.7845
2.7851
2.7804
2.7748
1.7B, 1M, 1×
2.6036
2.6017
2.5978
2.5941
2.5942
2.5922
1.7B, 4M, 4×
2.3859
2.3850
2.3814
2.3780
2.3751
2.3778
Appendix
Table 16: Final training loss for the cosine trajectory experiments. Bold marks the best method in each row.
Setup
Muon
Aurora
MuonEq
Muon 2
Aurora-LN
TensorChain
0.6B, 1M, 1×
2.8639
2.8624
2.8540
2.8497
2.8474
2.8379
0.6B, 4M, 4×
2.7527
2.7487
2.7434
2.7443
2.7391
2.7344
1.7B, 1M, 1×
2.5974
2.5964
2.5924
2.5885
2.5882
2.5857
1.7B, 4M, 4×
2.4823
2.4810
2.4784
2.4744
2.4734
2.4747
Appendix
Table 17: Final validation loss for the cosine trajectory experiments. Bold marks the best method in each row.
Figure 6: Cosine trajectories for Qwen3-0.6B, 1M-token batch, 1× Chinchilla.
Figure 7: Cosine trajectories for Qwen3-0.6B, 4M-token batch, 4× Chinchilla.
Figure 8: Cosine trajectories for Qwen3-1.7B, 1M-token batch, 1× Chinchilla.
Figure 9: Cosine trajectories for Qwen3-1.7B, 4M-token batch, 4× Chinchilla.
Table 18: TensorChain grouping ablation on Qwen3-0.6B, 1M-token batch, 1× Chinchilla. Each entry reports the best point from the corresponding full sweep.
Setting
Aurora-LN
Muon 2
MuonEq
Aurora
Muon
0.6B, 1M, 1×
8.2
9.0
10.3
12.4
12.8
0.6B, 4M, 4×
5.4
7.6
7.4
8.9
10.0
1.7B, 1M, 1×
3.9
4.0
6.2
7.5
7.8
1.7B, 4M, 4×
-2.0
-0.5
6.7
7.4
7.8
Appendix
Table 19: Cosine-schedule token-efficiency gain of TensorChain. Each entry is the percentage reduction in training tokens required to reach the same validation-loss target. Positive values favor TensorChain.