cs.LGNov 30, 2025

Provable Benefit of SignGD: A Minimal Model Under Heavy-Tailed Class Imbalance

Authors: Robin Yadav, Shuo Xie, Tianhao Wang, Zhiyuan Li

Organizations: Toyota Technological Institute at Chicago · University of British Columbia · University of California, San Diego

Abstract

Adaptive and non-Euclidean optimizers often outperform Euclidean methods such as stochastic gradient descent (SGD) in language modeling by a large margin. Existing theory usually explains this gap by assuming favorable smoothness geometry or noise structure tailored to the specific optimizer. We instead ask whether such geometry can be induced from a concrete learning setting. Starting from an optimizer gap that persists across realistic language-modeling experiments, we progressively remove sequence dependence, architectural complexity, and stochasticity. We find that the gap exists in a minimal setting: the softmax unigram model with heavy-tailed data. This model exposes a simple deterministic mechanism under heavy-tailed class imbalance. We prove that GD learns rare tokens slowly because the corresponding logits receive only tiny updates, while SignGD removes this magnitude dependence and moves rare and common coordinates on a more comparable scale. We make this precise with upper and lower bounds for the convergence rate of GD and upper bounds for the convergence of SignGD. Our stochastic bounds contain additional noise-dependent terms that can obscure this advantage in the convergence guarantees and can be reduced by increasing the batch size

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

    Jun 22, 2026Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2AdamLong-Tailed Distribution

  2. When and Why SignSGD Outperforms SGD: A Theoretical Study Based on ℓ1\ell_1-norm Lower Bounds

    May 7, 2026Hongyi Tao, Dingzhi Yu, Lijun ZhangStochastic Gradient DescentSignsgd