cs.LGMay 9, 2026

When and Why Grouping Attention Heads Accelerates Muon Optimization

Authors: Hongtao ZhangWenjie ZhouWei ChenXueqi Cheng

Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · University of Chinese Academy of Sciences

Abstract

Muon orthogonalizes matrix updates, but multi-head attention naturally operates at the level of heads. This granularity mismatch raises the question of whether Muon should be applied to the full attention projection, to individual heads, or to intermediate head groups. We study this question through a one-step descent comparison between full-matrix Muon and group-wise Muon. Our analysis reveals a trade-off between the \textbf{group-wise whitening gain} from group-wise updates and the \textbf{grouping-induced norm cost}, an additional update-norm cost caused by replacing full-matrix whitening with group-wise whitening. Motivated by this trade-off, we propose \textbf{Group Muon}, which treats head group size and grouping rule as optimizer hyperparameters. On GPT-2 Small trained on FineWeb, appropriate grouping improves validation loss over both full-QKV Muon and fully head-wise MuonSplit.

Explore similar work

CardsList
  1. Muown: Row-Norm Control for Muon Optimization

    May 11, 2026Kai Lion, Florian Hübler, Bingcong Li +2Muon OptimizerMuon

  2. Reassessing Muon for Matrix Factorization

    Jul 14, 2026Ali Parviz, Gal Mishne, Alex CloningerMuon OptimizerMuon

  3. Muown Implicitly Performs Angular Step-size Decay

    Jun 22, 2026Florian Hübler, Kai Lion, Antonio Orvieto +1Muon OptimizerMuon