cs.LGOct 7, 2026

Optimizing Large Language Models with Chained LMOs

Authors: Sungyoon Kim, Kaan Ozkara, Youngsuk Park

Organizations: Department of Electrical Engineering, Stanford University · Amazon Annapurna Labs

Abstract

Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining

    Sep 28, 2026Chang-Wei Shi, Xu Wang, Wu-Jun LiLanguage Model PretrainingEfficient Language Model Training

  2. ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training

    Oct 5, 2026Yuanshi Liu, Boyuan Jiang, Liang Hou +4Spectral RegularizationDeep Learning Optimization

  3. AMO: Operator-level Adaptive Muon Orthogonalization

    May 18, 2026Xinlin Zhuang, Panyi Ouyang, Yichen Li +7Language Model PretrainingNewton-Schulz Iteration