Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization
Authors: Junjie Xiao, Huiwen Jia
Organizations: Department of Mathematics, Peking University · Department of Industrial Engineering and Operations Research, University of California, Berkeley
Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch logistic gradient descent (GD) trajectories on strictly linearly separable data, with large initialization of scale R motivated by prior training. From any limiting normalized initial position, we use minimum-norm projection rules to construct a unique continuous ideal path consisting of finitely many linear segments. The path has two stages: negative-margin correction followed by minimum-margin growth. We prove that, after an explicit two-stage time reparameterization, the fixed-step GD trajectory divided by R converges uniformly to this path on every fixed parameter interval as R→∞. Further, our quantitative error bounds account for initialization perturbations and the transition between stages. This approximation provides asymptotic formulas for peak evaluation loss and cumulative training loss. In particular, peak evaluation loss can grow linearly in R even when both endpoint losses tend to zero. The cumulative losses in the correction and margin-growth stages, normalized by R2 and R, respectively, converge to explicit limits. Experiments on controlled geometries and fixed image features complement our theoretical results.
Figures & tables
Figure 1: The two-stage ideal path. The analytic example uses z1=(2,0) , z2=(1,2) , p1=p2=1/2 , and u0=(−1,−1/5) . (a) The dashed correction path turns when the second margin reaches zero and reaches the shaded nonnegative region C at E=(0,1/5) . The solid continuation follows z1 until the margins meet, then follows their equal-margin line with velocity (8/5,4/5) . (b) The same path’s margins, zi⊤P(r) , show the correction turn at r=2/5 , entry at τ∗=4/5 , and persistent equality from r=1 . The horizontal axis uses the two-stage parameter defined in ( 5 ).
Figure 2: An intermediate peak in evaluation loss. The analytic example follows the training geometry of Figure 1 in the invariant plane x3=1 . (a) The path crosses the shaded evaluation region a⊤x<0 ; correction is dashed and negative evaluation margins are marked in orange. (b) The limiting normalized evaluation loss peaks at 1/25 , although it vanishes at both endpoints. Both panels are computed from the ideal path.
Figure 3: GD approaches the complete ideal path. (a) A coordinate projection of the four-sample example, omitting most of the initial approach to highlight the turns. Colored curves connect saved GD iterates divided by R ; the dashed line shows the ideal path, with its entry E and endpoint P(S) marked. (b) Uniform path errors for the analytic and four-sample examples. (c) The median and 10th–90th percentile band over all 40 random instances. Errors in (b,c) use full parameter vectors over each problem’s complete window [0,S] .
Figure 4: Functional predictions and a test on image features. Panels (a,b) use the four-sample example. (a) Evaluation loss divided by R , compared with Hev(P(r)) . Curves connect saved GD iterates; shading marks ideal correction. (b) Cumulative training losses divided by 0.38277R2 before entry and 0.10131R after entry; the dashed line at one is the limit. (c) Uniform path error for the 32-image digit task over its fixed window.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Runs
Number
Total seconds
Largest update count
Budget per run
G1, short window
5
0.63
5.395×106
4×109
G2, main window
5
8.09
5.962×107
4×109
Random instances
200
187.93
1.415×109
2×109
32-image digits task
5
452.75
2.691×107
2×108
Larger digits tasks, entry
12
382.82
4.785×106
5×107
Appendix
Table 1: Recorded computation for the selected numerical results. The last column is the allowed number of updates per run; every run in this table completed its specified stopping rule.
j
tj
tj+1
vj
0
0
4/15
(−3/4,7/4,5/4)
1
4/15
121/360
(−1/4,5/4,3/2)
2
121/360
151/360
(1/2,1/2,5/4)
3
151/360
383/864
(2/9,7/9,10/9)
4
383/864
28369/54432
(1/2,1/4,3/4)
5
28369/54432
85637/163296
(2,1,3)
Appendix
Table 2: Finite construction of G2 on the reported window. All event times and velocities shown as fractions are exact; the final time is the original grid-selected S .
Instance
η
S
Instance
η
S
00
2.383566011883
0.812550874730
20
3.239836363412
0.687360271366
01
4.988250994041
3.192085672009
21
3.359344012717
1.983970959559
02
4.079246859906
3.703531488345
22
2.326955393044
2.986875270919
03
3.378190352060
2.424884905424
23
3.180895784073
0.292845536333
04
1.769894663824
0.691047805794
24
3.970855233719
1.771797801208
05
1.661883647012
2.979381305143
25
3.614814813406
0.875733263005
Appendix
Table 3: Step sizes and fixed windows for all random instances, rounded to twelve decimal places. The generation sequence and grid rule specify their full floating-point values.
R
After entry
Entry at stop
No entry
10%
Median
90%
5
14
20
6
0.07693
0.18467
0.36983
10
20
15
5
0.05373
0.13004
0.21629
20
26
11
3
0.04282
0.09769
0.15039
40
36
3
1
0.03021
0.05787
0.09460
80
39
1
0
0.01850
0.03701
0.06194
Appendix
Table 4: All 40 random runs complete the fixed window at each value of R . “After entry” counts runs with at least one post-entry update; the next two columns give entry at the stopping node and no entry by that node.
Geometry
η
f(R)
ηΛ
ER(S)
Entry error
G1
0.25
0
0.20505
0.00795
0.00528
G1
0.5
0
0.41010
0.01160
0.00664
G1
1
0
0.82019
0.00885
0.00799
G1
2
0
1.64039
0.01433
0.01383
G1
4
0
3.28078
0.02796
0.02790
G1
1
1
0.82019
0.01398
0.00945
Appendix
Table 5: Completed stability checks at R=80 . The entry error is ∥wκR/R−E∥ relative to the unperturbed ideal entry.
Source
Target
Source size
Train/test
Source margin
η
τ∗
3/8
5/6
357
254/109
0.20773
0.328481301706
6.473229321622
4/9
2/3
361
251/109
0.37357
0.317649989551
14.787703491645
0/1
7/9
360
251/108
0.59168
0.342935275073
16.407770420857
Appendix
Table 6: Larger digits tasks used for the entry-location comparison. The source margin is miniziA⊤u0 . The target split and feature map are specified above.
Figure 5: Additional functional and entry-location results. (a) The maximum evaluation loss divided by R in the full-rank four-sample geometry G2 approaches the ideal value 553/(144017) ; maxima include every GD iterate through the stopping iterate. (b) Mean loss on all 357 source images of digits 3/8 , divided by R , along GD trained on the fixed 32 images of digits 5/6 . The black curve is the ideal evaluation profile, and shading marks the ideal correction interval. Colored curves connect saved nodes using each run’s actual time parameterization. (c) Entry-position errors ∥wκR/R−E∥ for the three larger target tasks in Table 6 . Each run stops at actual entry; this panel measures entry convergence rather than a completed post-entry path.