math.OCJun 14, 2026

Schattor: Schatten-family methods for deep learning optimization

Authors: Bohao Ma, Junyu Zhang, Chuan He

Organizations: School of Data Science, The Chinese University of Hong Kong (Shenzhen), Shenzhen, Guangdong, China · Department of Industrial Systems Engineering and Management, National University of Singapore, Singapore · Department of Mathematics, Linköping University, Sweden

Abstract

Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis. Motivated by the limitations of SGD and the success of adaptive optimizers, we propose {\it Schattor}, a family of adaptive first-order methods based on Schatten norms. Schattor unifies SGD and the recently proposed matrix-variate adaptive optimizer Muon within a single Schatten-norm-based framework. We establish dimension-free stationarity guarantees for methods in the Schattor family for stochastic matrix optimization problems via a novel matrix martingale moment bound. We also develop multi-block extensions that adaptively balance block-wise optimization progress and prove dimension-free stationarity guarantees in this more general setting.

Explore similar work

CardsList
  1. From SGD to Muon: Adaptive Optimization via Schatten-p Norms

    May 19, 2026Thomas Massena, Corentin Friedrich, Mathieu SerrurierMuon OptimizerOptimizer Design