Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.
Figures & tables
Figure 1.1: LLaMA 130M training dynamics. Effective rank is measured on a middle-layer V projection; the dashed line marks regularization release.
Figure 2.1: Single-layer V -spectrum intervention with O -only Muon updates. GO and UO denote the gradient and Muon update direction. In (c), rollout steps and loss decrease use logarithmic scales.
Figure 2.2: Validation loss under persistent spectral constraints.
Figure 4.1: LLaMA pretraining results.
Figure 4.2: Token efficiency and shaping ablations in LLaMA Pretraining.
Optimizer
ARC-C
ARC-E
HellaSwag
PIQA
WinoGrande
Avg.
Muon
0.305(0.013)
0.342(0.010)
0.549(0.005)
0.705(0.011)
0.547(0.014)
0.489
ORCA
0.311(0.014)
0.357(0.010)
0.573(0.005)
0.711(0.011)
0.553(0.014)
0.501
Table 4.1: Downstream evaluation on LLaMA 1.3B.
Figure 4.3: Validation loss across architectures and base optimizers.
Figure 4.4: Fine-grained MoE validation loss with layer-wise release.
Figure 5.1: Loss curves under different shaping windows on LLaMA 130M.
Figure 5.2: Alternative early spectral interventions on LLaMA 130M, released to Muon after 25% of training.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Layers
Hidden
FFN
Q/KV heads
Head dim
130M
12
768
2048
12/12
64
350M
24
1024
2736
16/16
64
1.3B
24
2048
5461
32/32
64
Appendix
Table B.1: LLaMA configurations used in the standard experiments.
Model
Layers
Hidden
Dense FFN
Q/KV heads
Experts
Top- k
Qwen3 160M
20
512
1536
8/4
–
–
Qwen3 0.6B
28
1024
3072
16/8
–
–
MoE 0.8B
16
768
1536
6/3
64
4
MoE 8B
27
2048
4096
16/8
64
4
Appendix
Table B.2: Qwen3 and fine-grained MoE configurations.
Model
Data
Seq. len.
Global batch
Steps
Tokens
Peak LR
WD
LLaMA 130M
C4
1024
512
50,000
26.2144B
searched
0.01
LLaMA 350M
C4
1024
1024
100,000
104.8576B
1×10−3
0.1
LLaMA 1.3B
C4
1024
1024
100,000
104.8576B
6×10−4
0.1
Qwen3 160M
The Pile
4096
256
25,000
26.2144B
1×10−3
0.01
Qwen3 0.6B
The Pile
4096
512
50,000
104.8576B
6×10−4
0.1
MoE 0.8B
olmo-mix-1124
1024
512
50,000
26.2144B
1×10−3
0.1
Appendix
Table B.3: Training budgets for the standard experiments. Tokens are computed as sequence length times global batch size times optimizer steps.
Method
Seed 0
Seed 1
Seed 2
Seed 3
Mean
Variance
Adam
2.882
2.880
2.877
2.884
2.8808
8.92×10−6
Muon
2.846
2.851
2.852
2.853
2.8505
9.67×10−6
ORCA
2.811
2.813
2.812
2.811
2.8118
9.17×10−7
Appendix
Table C.1: Final validation loss across four random seeds for LLaMA 130M at 50k steps. The final two columns report the mean and sample variance across seeds; summary statistics are computed from the displayed values.
Question
Reference setting
Variant setting
Reference loss ↓
Variant loss ↓
Variant − Reference
Optimizer state
ORCA, state retained
ORCA, state reset
2.8107
2.8106
−0.0001
Schedule: WSD
Muon
ORCA
2.8512
2.8106
−0.0406
Schedule: cosine
Muon
ORCA
2.8595
2.8139
−0.0456
Appendix
Table C.2: Optimizer-state and learning-rate-schedule controls.
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iteration as its core operation. Standard Muon applies a uniform NS schedule to all parameter matrices, overlooking possible differences in orthogonalization difficulty and its impact on performance. Through a systematic empirical study, we show that this per-matrix heterogeneity is pervasive and strongly associated with matrix geometry, which evolves dynamically across operator types, training stages, and network depths. Therefore, uniform NS schedules can lead to uneven orthogonalization quality across the model. Motivated by these findings, we propose Operator-level Adaptive Muon Orthogonalization (AMO), an observe-then-commit method that measures weight geometry by operator type early in training and then uses these signals to allocate the NS budget for the remainder of training. AMO delivers consistent improvements over uniform-schedule Muon across standard, prolonged, and continual pre-training, surpassing the strongest baseline by +0.76 on Llama3.1-1.4B and +0.51 on Qwen3-1.7B in average downstream performance of 12 evaluation tasks, with gains persisting at Llama3.1-4B scale.
Xinlin Zhuang, Panyi Ouyang, Yichen Li +7
The Chinese University of Hong Kong · Shopee · MBZUAI +3
Low-rank adaptation (LoRA) makes finetuning large language models cheaper by adding to each weight matrix a trainable low-rank update parameterized as the product of two matrices. These matrices are usually trained with Adam, which treats them as a single flat vector of parameters and ignores both the matrix and product structure of LoRA. Applying a matrix-aware optimizer such as Muon to each factor does not consistently improve over Adam, and neither do the product-aware Muon variants proposed in concurrent works. To realize consistent gains, we introduce PoLoRA, a Preconditioned Orthogonalized LoRA optimizer built from three ingredients: a product-aware spectral update direction, curvature preconditioning derived from controlling the per-sample loss change, and a magnitude rule that controls the sizes of both the factor and merged updates. We evaluate PoLoRA on instruction-tuning datasets for code and math across models from 1B to 8B parameters, and find that it reaches the final held-out loss achieved by tuned Adam in 1.2-1.7 times fewer steps, while adding at most 3% per-step overhead. Compared to Adam, PoLoRA is also less sensitive to the learning rate, and its optimal learning rate is stable across ranks.
Nikhil Ghosh, Tetiana Parshakova, Robert M. Gower
Center for Computational Mathematics, Flatiron Institute, New York, NY
The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term introduces curvature-sensitive corrections while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.