cs.LGOct 2, 2026

Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization

Authors: Junjie Xiao, Huiwen Jia

Organizations: Department of Mathematics, Peking University · Department of Industrial Engineering and Operations Research, University of California, Berkeley

Abstract

Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch logistic gradient descent (GD) trajectories on strictly linearly separable data, with large initialization of scale RR motivated by prior training. From any limiting normalized initial position, we use minimum-norm projection rules to construct a unique continuous ideal path consisting of finitely many linear segments. The path has two stages: negative-margin correction followed by minimum-margin growth. We prove that, after an explicit two-stage time reparameterization, the fixed-step GD trajectory divided by RR converges uniformly to this path on every fixed parameter interval as R→∞R\to\infty. Further, our quantitative error bounds account for initialization perturbations and the transition between stages. This approximation provides asymptotic formulas for peak evaluation loss and cumulative training loss. In particular, peak evaluation loss can grow linearly in RR even when both endpoint losses tend to zero. The cumulative losses in the correction and margin-growth stages, normalized by R2R^2 and RR, respectively, converge to explicit limits. Experiments on controlled geometries and fixed image features complement our theoretical results.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics

    Aug 5, 2026Han BaoGradient DescentOverparameterization

  2. Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes

    Jun 7, 2024Si Yi Meng, Antonio Orvieto, Daniel Yiming Cao +1Step AccuracyGradient Descent

  3. Improved Convergence of Large Stepsize Gradient Descent for Logistic Regression

    Oct 5, 2026Xiaochuan Gong, Ang LiLogistic Regression