cs.LGOct 5, 2026

ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training

Authors: Yuanshi Liu, Boyuan Jiang, Liang Hou, Xin Tao, Pengfei Wan, Zhouchen Lin, Cong Fang

Organizations: School of Intelligence Science and Technology, Peking University · Kling Team

Abstract

Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AMO: Operator-level Adaptive Muon Orthogonalization

    May 18, 2026Xinlin Zhuang, Panyi Ouyang, Yichen Li +7MuonOrthogonality

  2. PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer

    Jul 20, 2026Nikhil Ghosh, Tetiana Parshakova, Robert M. GowerMultimodal Continual Instruction TuningMuon