Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.
Figures & tables
Figure 1.1: LLaMA 130M training dynamics. Effective rank is measured on a middle-layer V projection; the dashed line marks regularization release.
Figure 2.1: Single-layer V -spectrum intervention with O -only Muon updates. GO and UO denote the gradient and Muon update direction. In (c), rollout steps and loss decrease use logarithmic scales.
Figure 2.2: Validation loss under persistent spectral constraints.
Figure 4.1: LLaMA pretraining results.
Figure 4.2: Token efficiency and shaping ablations in LLaMA Pretraining.
Optimizer
ARC-C
ARC-E
HellaSwag
PIQA
WinoGrande
Avg.
Muon
0.305(0.013)
0.342(0.010)
0.549(0.005)
0.705(0.011)
0.547(0.014)
0.489
ORCA
0.311(0.014)
0.357(0.010)
0.573(0.005)
0.711(0.011)
0.553(0.014)
0.501
Table 4.1: Downstream evaluation on LLaMA 1.3B.
Figure 4.3: Validation loss across architectures and base optimizers.
Figure 4.4: Fine-grained MoE validation loss with layer-wise release.
Figure 5.1: Loss curves under different shaping windows on LLaMA 130M.
Figure 5.2: Alternative early spectral interventions on LLaMA 130M, released to Muon after 25% of training.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Layers
Hidden
FFN
Q/KV heads
Head dim
130M
12
768
2048
12/12
64
350M
24
1024
2736
16/16
64
1.3B
24
2048
5461
32/32
64
Appendix
Table B.1: LLaMA configurations used in the standard experiments.
Model
Layers
Hidden
Dense FFN
Q/KV heads
Experts
Top- k
Qwen3 160M
20
512
1536
8/4
–
–
Qwen3 0.6B
28
1024
3072
16/8
–
–
MoE 0.8B
16
768
1536
6/3
64
4
MoE 8B
27
2048
4096
16/8
64
4
Appendix
Table B.2: Qwen3 and fine-grained MoE configurations.
Model
Data
Seq. len.
Global batch
Steps
Tokens
Peak LR
WD
LLaMA 130M
C4
1024
512
50,000
26.2144B
searched
0.01
LLaMA 350M
C4
1024
1024
100,000
104.8576B
1×10−3
0.1
LLaMA 1.3B
C4
1024
1024
100,000
104.8576B
6×10−4
0.1
Qwen3 160M
The Pile
4096
256
25,000
26.2144B
1×10−3
0.01
Qwen3 0.6B
The Pile
4096
512
50,000
104.8576B
6×10−4
0.1
MoE 0.8B
olmo-mix-1124
1024
512
50,000
26.2144B
1×10−3
0.1
Appendix
Table B.3: Training budgets for the standard experiments. Tokens are computed as sequence length times global batch size times optimizer steps.
Method
Seed 0
Seed 1
Seed 2
Seed 3
Mean
Variance
Adam
2.882
2.880
2.877
2.884
2.8808
8.92×10−6
Muon
2.846
2.851
2.852
2.853
2.8505
9.67×10−6
ORCA
2.811
2.813
2.812
2.811
2.8118
9.17×10−7
Appendix
Table C.1: Final validation loss across four random seeds for LLaMA 130M at 50k steps. The final two columns report the mean and sample variance across seeds; summary statistics are computed from the displayed values.
Question
Reference setting
Variant setting
Reference loss ↓
Variant loss ↓
Variant − Reference
Optimizer state
ORCA, state retained
ORCA, state reset
2.8107
2.8106
−0.0001
Schedule: WSD
Muon
ORCA
2.8512
2.8106
−0.0406
Schedule: cosine
Muon
ORCA
2.8595
2.8139
−0.0456
Appendix
Table C.2: Optimizer-state and learning-rate-schedule controls.