cs.LGSep 27, 2026

On the Two Faces of Adam in Separable Linear Classification

Authors: Chen Fan, Csaba Szepesvári

Organizations: University of Alberta

Abstract

We consider the behavior of deterministic, full-batch, bias-corrected Adam in separable linear classification with softmax parametrization under log-loss. In this setting, under a wide range of conditions Adam is known to approach max-norm-margin optimality when its stability constant εε is zero, while with a positive εε, it is known to approach Euclidean-margin optimality. Our main contribution is the quantitative description of Adam's behavior for small fixed positive εε. We give sufficient conditions under which an Adam-trained classifier nearly maximizes the max-norm margin before the updates become gradient-like. We also show that the classifier reaches a fixed target Euclidean margin only much later. Specifically, we show that for polynomially decreasing stepsizes with exponent aa, where 1/3<a<11/3<a<1, the updates become approximately proportional to the negative gradient after Θ(log⁡(1/ε)1/(1−a))Θ(\log(1/ε)^{1/(1-a)}) iterations. At that time, the classifier still nearly maximizes the max-norm margin. Reaching a fixed target Euclidean margin above that of every max-norm-optimal classifier, but below the optimum, is shown to require ε−Θ(1)/(1−a)ε^{-Θ(1)/(1-a)} iterations. Under inverse-linear stepsize decay (a=1a=1), the update transition takes polynomially many iterations, whereas reaching the target margin takes exponentially many. Experiments support these predictions. The later change in the classifier can improve or worsen generalization after training error reaches zero, connecting the analysis to grokking and its reverse.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Stability Annealing Selects the Implicit Bias of Smoothed Sign Descent: A Rate-Indexed Barrier Path on Separable Data

    Jul 7, 2026Xiangwu Wang, Chengwei Cao, Yicheng Song +2Mirror DescentAnnealing

  2. Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization

    Date pendingSharan Sahu, Abir Sarkar, Cameron J. Hogan +1AdamStationarity

  3. Adam at the Edge of Stability: Adaptive Feedback, Provable Oscillation, and Gradient Reversal

    Aug 21, 2026Yiman Fong, Heng YangAdamLocal Curvature