cs.LGSep 28, 2026

MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining

Authors: Chang-Wei Shi, Xu Wang, Wu-Jun Li

Organizations: National Key Laboratory for Novel Software Technology, School of Computer Science, Nanjing University, P. R. China

Abstract

The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Muown: Row-Norm Control for Muon Optimization

    May 11, 2026Kai Lion, Florian Hübler, Bingcong Li +2MuonOptimizer Design

  2. AMO: Operator-level Adaptive Muon Orthogonalization

    May 18, 2026Xinlin Zhuang, Panyi Ouyang, Yichen Li +7MuonOrthogonality

  3. Variance-Adaptive Muon: Pre-Orthogonalization Variance Modulation for Efficient Language Model Pretraining

    Jan 21, 2026Jingru Li, Yibo Fan, Huan LiMuonLarge Language Model Pretraining