We study logistic regression on linearly separable data under gradient descent with a large constant stepsize η. Such dynamics may exhibit a characteristic Edge of Stability phenomenon, in which the loss initially oscillates before transitioning to a stable phase of monotone decrease. Existing work provides a tight Θ(1) bound in dimension d=2 as η→∞ and conjectures a bound independent of η in arbitrary dimensions d≥2. In this paper, we disprove this conjecture by showing that, for every fixed sample size n≥2 and sufficiently small margin γ, the worst-case transition time is Θ((logη)min{n−2,d−2}) uniformly over d≥2. The key challenge in establishing a tight bound is that the sample contributing most strongly to the gradient can change repeatedly across iterations. To address this issue, we control such changes by induction on dimension and sample size, and construct matching hard instances.
Figures & tables
Figure 1: The transition time can grow logarithmically with the stepsize when n=d=3 . (a) Both transition times ση(D) and τη(D) grow approximately linearly with logη , with a separate dataset for each stepsize and a common positive margin lower bound. (b,c) The loss continues to oscillate after all samples are correctly classified, as the two interacting samples alternate in driving the updates. Shading marks the stable phase. See for the details of the simulation.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 2: The total loss can keep increasing while the dominant sample switches. For the symmetric construction with n=d=3 , (a) both transition times grow approximately linearly with logη , with a separate dataset for each stepsize. (b,c) Even after all samples are correctly classified, the loss keeps increasing as the two interacting samples alternate in driving the updates. Panel (c) shows their relative update coefficients. Shading marks the stable phase.