cs.LGJun 13, 2026

When to use what Schatten-pp norm in deep learning?

Authors: Thomas Pethick

Abstract

Schatten-\infty based optimizers such as Muon have shown promising empirical performance, but there remains seemingly conflicting observations regarding whether they are beneficial. We resolve this conflict by showing that the conclusion is regime dependent. Even when the objective is smooth in the Schatten-\infty geometry, smaller Schatten-pp geometries can be optimal, specifically in the low-dimensional regime, which we show includes Chinchilla scaling. This conclusion follows from a new noise-robust acceleration result for the SODA framework for p>2p>2. The same analysis explains why Muon-like methods do not require warmup, why they naturally favor large batches, and yields a batch size scaling rule for arbitrary pp.

Explore similar work

CardsList
  1. Muon is Not That Special: Random or Inverted Spectra Work Just as Well

    May 11, 2026Zakhar Shumaylov, Nathaël Da Costa, Peter Zaika +6Muon OptimizerMuon