cs.LGJan 29, 2026

Why β1=β2β_1 = β_2 Is Dynamically Special in Adam

Authors: Alberto Fernández-Hernández, Cristian Pérez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ortí

Organizations: Universitat Politècnica de València Valencia, Spain · Universitat Jaume I Castelló de la Plana, Spain

Abstract

Adam has been at the core of large-scale training for almost a decade, yet the role of its two momentum parameters remains poorly understood. Recent work shows that tying β1=β2β_{1}=β_{2} can preserve Adam's strong performance despite collapsing two memory scales into one, raising a basic question: what becomes dynamically special when the memories are tied? We identify a concrete mechanism. In the continuous-time limit, each normalized-update coordinate decomposes into a sign component, an explicit magnitude-lag term proportional to the difference between the two memory times, and additional transition, curvature, and nonlinear ratio terms. This lag channel vanishes exactly when β1=β2β_{1}=β_{2}, making the diagonal the unique regime in which this mismatch-induced response is structurally absent. A full-history discrete decomposition on real training gradients recovers this change in composition: tied updates are sign-dominated, whereas the lag term becomes substantial off the diagonal and leaves a comparatively small residual. Across six vision and language tasks, tied configurations also typically exhibit smoother update-norm trajectories. Overall, our results identify memory-scale mismatch as a concrete source of magnitude sensitivity in Adam and provide a mechanistic account of why tied momentum is dynamically distinctive.

Explore similar work

CardsList
  1. Refresh-Scaling the Memory of Balanced Adam

    May 11, 2026Alberto Fernández-Hernández, Cristian Pérez-Corral, Jose I. Mestre +2AdamBalanced Learning

  2. Beyond Quadratic Loss: The Stability Phase Diagram of Adam

    Sep 16, 2026Gaoxiang Tang, Huanran Chen, Ziming LiuAdamLoss Landscape

  3. Adaptive Momentum and Nonlinear Damping for Neural Network Training

    Jan 30, 2026Aikaterini Karoni, Rajit Rajpal, Benedict Leimkuhler +1Adam