Organizations: Department of Mathematics, University of California, Los Angeles · School of Data, Mathematical, and Statistical Sciences, University of Central Florida
Interior-point methods (IPMs) are among the most widely used algorithms for constrained optimization, yet their Newton-based search directions require costly second-order information and large linear-system solves. Learning to optimize offers cheaper updates learned from data, but the singular behavior of logarithmic barriers near constraint boundaries makes IPMs highly sensitive to perturbations, complicating both warm starting and learning reliable updates. We introduce pdLIP, an IPM for smooth nonlinear programs that integrates learned preconditioning with pdProj, an all-shifted primal-dual projected-search IPM. A shared coordinate-wise recurrent network predicts a positive diagonal preconditioner that scales the right-hand side of the reduced Newton system for the primal step, and the remaining slack and multiplier directions are recovered analytically. The learned iterations avoid Hessian evaluations and Newton-system solves, using only first-order and coordinate-wise operations amenable to GPU parallelization. Training is self-supervised, with a loss based on a penalty-barrier merit function and the residual of perturbed optimality conditions, requiring neither target directions nor precomputed solutions. Primal and dual shifts mitigate the barrier's sensitivity to perturbations near constraint boundaries, enabling effective warm starting. Across four classes of 200-dimensional convex and nonconvex constrained problems, pdLIP warm starts reduce pdProj refinement iterations by 63-67% compared with cold starts at the same KKT residual tolerance of 10−8, with negligible warm-start generation cost relative to the subsequent pdProj solve. Improvements persist on box-constrained QPs with 1000 variables and extend to applications including portfolio optimization, support vector machines, and a nonlinear control example.
Figures & tables
Figure 1: Overview of one pdLIP iteration. At iteration k , the shared LSTM–MLP maps each ϕk,j to a raw scaling pk,j . The transformation pk,j=∣pk,j∣+δ yields a positive diagonal scaling Pk≻0 , which defines the primal direction Δxk=−Pkqk . The remaining components of Δvk are recovered analytically from the pdProj equations, after which the full trial step is projected onto Ωk . The estimates and barrier and penalty parameters are then updated based on pdProj O/M/F rules.
Class
IPOPT
Cold pdProj
pdLIP + pdProj
Reduction (Iter. / Time)
Iter.
Iter. / Time
WS Cost
Iter. / Time
Total Time
pdLIP
IPM-LSTM
C-RHS
15.63
13.41/7.53 s
20.11 ms
4.64/2.58 s
2.60 s
65.40/65.48%
42.4/15.1%
C-ALL
15.82
13.22/7.41 s
19.93 ms
4.35/2.42 s
2.44 s
67.10/67.11%
35.7/7.5%
NC-RHS
15.60
13.47/7.53 s
20.03 ms
4.93/2.78 s
2.80 s
63.38/62.81%
27.5/1.3%
NC-ALL
15.72
13.46/8.85 s
20.20 ms
4.84/2.79 s
2.81 s
64.04/68.29%
15.4/−8.6%
Table 1: Warm-start results on the constrained benchmark problems from ( Gao et al., 2024 ) .
Dim.
IPOPT
Cold pdProj
pdLIP + pdProj
Reduction
Iter.
Iter. / Time
WS Cost
Iter. / Time
Total Time
Iter. / Time
n=10
8.99
7.45/6.66 ms
0.98 ms
1.52/1.80 ms
2.78 ms
79.55/58.21%
n=200
13.48
11.44/6.62 s
7.32 ms
2.19/1.17 s
1.18 s
80.81/82.17%
n=1000
15.50
13.87/13.74 s
40.94 ms
7.23/4.53 s
4.57 s
47.90/66.73%
Table 2: Warm-start results on convex box-constrained QPs across dimensions.
Problem
(s,t)
IPOPT
Cold pdProj
pdLIP + pdProj
Reduction
Iter.
Iter. / Time
WS Cost
Iter. / Time
Total Time
Iter. / Time
Portfolio
(50,5)
14.68
8.99/22.94 ms
5.16 ms
3.48/7.82 ms
12.98 ms
61.28/43.43%
Portfolio
(200,20)
21.06
12.07/6.50 s
14.07 ms
9.36/5.14 s
5.15 s
22.40/20.71%
SVM
(5,50)
12.52
9.30/3.45 s
11.57 ms
4.52/1.66 s
1.67 s
51.39/51.55%
SVM
(20,200)
15.46
13.83/10.21 s
35.95 ms
6.31/3.88 s
3.92 s
54.42/61.65%
Table 3: Warm-start results on portfolio optimization and SVM problems.
Problem
IPOPT
Cold pdProj
pdLIP + pdProj
Reduction
Iter.
Iter. / Time
WS Cost
Iter. / Time
Total Time
Iter. / Time
Quadrotor
9.20
9.74/19.05 s
71.99 ms
8.00/15.74 s
15.81 s
17.89/17.00%
Table 4: Warm-start results on the 200 -dimensional quadrotor optimization problem.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
10−4
104
0.0071
0.0027
0.0076
2
0.1
103
0.0013
0.0090
0.0091
3
0.1
5×103
0.0134
0.0024
0.0136
4
10−3
103
0.0092
0.0133
0.0162
5
0.01
103
0.0074
0.0211
0.0223
Appendix
Table 5: Hyperparameter sensitivity for convex QP RHS, n=50,meq=25,mineq=25 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
0.01
5×103
0.0013
0.0018
0.0022
2
10−3
5×103
0.0328
0.0034
0.0330
3
10−3
103
0.0335
0.0373
0.0501
4
0.1
5×103
0.0487
0.0142
0.0507
5
0.01
103
0.1082
0.0278
0.1117
Appendix
Table 6: Hyperparameter sensitivity for convex QP ALL, n=50,meq=25,mineq=25 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
0.1
5×103
0.0072
0.0030
0.0078
2
0.1
103
0.0014
0.0086
0.0087
3
10−3
103
0.0083
0.0091
0.0123
4
0.01
5×103
0.0326
0.0035
0.0327
5
0.01
104
0.0507
0.0065
0.0511
Appendix
Table 7: Hyperparameter sensitivity for nonconvex QP RHS, n=50,meq=25,mineq=25 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
10−3
103
0.0062
0.0119
0.0134
2
10−3
5×103
0.0119
0.0079
0.0143
3
0.1
5×103
0.0135
0.0074
0.0154
4
0.01
5×103
0.0154
0.0073
0.0171
5
0.1
104
0.0183
0.0073
0.0197
Appendix
Table 8: Hyperparameter sensitivity for nonconvex QP ALL, n=50,meq=25,mineq=25 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
0.01
104
0.0084
0.0061
0.0103
2
10−3
104
0.0112
0.0062
0.0128
3
0.01
5×103
0.0101
0.0101
0.0142
4
10−3
103
0.0198
0.0368
0.0418
5
0.1
104
0.0471
0.0068
0.0476
Appendix
Table 9: Hyperparameter sensitivity for convex QP RHS, n=200,meq=100,mineq=100 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
10−3
104
0.0057
0.0057
0.0080
2
0.01
104
0.0069
0.0059
0.0091
3
10−3
5×103
0.0042
0.0096
0.0105
4
0.01
5×103
0.0208
0.0106
0.0233
5
0.01
103
0.0053
0.0350
0.0354
Appendix
Table 10: Hyperparameter sensitivity for convex QP ALL, n=200,meq=100,mineq=100 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
0.01
5×103
0.0211
0.0152
0.0260
2
0.1
5×103
0.0203
0.0166
0.0262
3
0.1
104
0.0340
0.0110
0.0358
4
10−3
5×103
0.0502
0.0186
0.0535
5
10−3
103
0.0205
0.0512
0.0551
Appendix
Table 11: Hyperparameter sensitivity for nonconvex QP RHS, n=200,meq=100,mineq=100 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
0.01
5×103
0.0214
0.0236
0.0318
2
10−3
5×103
0.0303
0.0220
0.0375
3
0.01
104
0.0400
0.0138
0.0423
4
0.1
104
0.0408
0.0152
0.0436
5
10−3
103
0.0152
0.0508
0.0530
Appendix
Table 12: Hyperparameter sensitivity for nonconvex QP ALL, n=200,meq=100,mineq=100 .
Rank
μB0
μP0
τM
LR
Mean eP
Mean eD
ecomb
1
0.01
0.5
1
10−4
0.0019
0.0092
0.0094
2
0.01
2
1
10−4
0.0035
0.0110
0.0116
3
0.01
1
1
10−4
0.0038
0.0120
0.0126
4
0.01
2
5
10−4
0.0039
0.0134
0.0139
5
0.01
1
5
10−4
0.0092
0.0225
0.0243
Appendix
Table 13: Hyperparameter sensitivity for portfolio problems, s=50,t=5 .
Rank
μB0
μP0
τM
LR
Mean eP
Mean eD
Combined
1
0.01
0.1
1
10−4
0.0091
0.0988
0.0992
2
0.01
0.5
1
10−4
0.0101
0.1170
0.1175
3
10−3
5×103
1
10−4
0.0569
0.1688
0.1782
4
0.01
2
1
10−4
0.0135
0.1800
0.1805
5
0.01
1
1
10−4
0.0074
0.2256
0.2257
Appendix
Table 14: Hyperparameter sensitivity for portfolio problems, s=200,t=20 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
10−3
0.5
0.0003
0.0086
0.0086
2
0.1
0.5
0.0005
0.0200
0.0200
3
0.01
2
0.0005
0.0312
0.0312
4
0.1
2
0.0006
0.0761
0.0761
5
10−3
2
0.0023
0.0969
0.0969
Appendix
Table 15: Hyperparameter sensitivity for SVM problems, s=5,t=50 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
10−3
0.5
0.0002
0.0324
0.0324
2
0.01
0.5
0.0007
0.1345
0.1345
3
10−3
5
0.0002
0.1726
0.1726
4
0.1
2
0.0039
0.1805
0.1805
5
10−3
2
0.0065
0.3435
0.3436
Appendix
Table 16: Hyperparameter sensitivity for SVM problems, s=20,t=200 .
Rank
μB0
μP0
Mean eP
Mean eD
ecomb
1
10−2
0.1
4.0483×10−5
0.0304
0.0304
2
10−2
2
0.0096
0.0414
0.0426
3
10−3
2
0.0082
0.0460
0.0467
4
10−3
10
0.0299
0.0811
0.0865
5
10−2
10
0.0352
0.1168
0.1220
Appendix
Table 17: Hyperparameter sensitivity for single-quadrotor navigation, n=200 and T=11 .
Problem Class
Method
Iter.
Grad. Calls
Func. Calls
Convex QP RHS
IPOPT
15.63
17.63
16.63
Cold pdProj
13.41
16.73
16.73
Warm-started pdProj
4.64
5.64
5.64
Convex QP ALL
IPOPT
15.82
17.82
16.82
Cold pdProj
13.22
17.65
17.65
Warm-started pdProj
4.35
5.36
5.36
Appendix
Table 18: Iterations, gradient calls, and function calls for the n=200 constrained benchmarks.
Problem Class
Method
Iter.
Grad. Calls
Func. Calls
Convex QP RHS
IPOPT
12.08
14.08
13.08
Cold pdProj
8.97
10.38
10.38
Warm-started pdProj
2.92
3.93
3.93
Iteration reduction
67.49%
–
–
Convex QP ALL
IPOPT
12.17
14.17
13.17
Cold pdProj
9.31
11.00
11.00
Appendix
Table 19: Iterations, gradient calls, and function calls for the n=50 constrained benchmarks.
Problem Class
Cold Time
WS Cost
Total Time
Time Red.
Convex QP RHS
2.69 s
5.23 ms
878 ms
67.31%
Convex QP ALL
2.80 s
4.59 ms
861 ms
69.22%
Nonconvex RHS
2.81 s
5.31 ms
886 ms
68.44%
Nonconvex ALL
2.84 s
5.64 ms
917 ms
67.75%
Appendix
Table 20: Runtime results for the n=50 constrained benchmarks.
Problem
Method
Iter.
Grad. Calls
Func. Calls
n=200
IPOPT
13.48
15.48
14.48
Cold pdProj
11.44
12.44
12.44
Warm-started pdProj
2.19
3.19
3.19
n=1000
IPOPT
15.50
17.50
16.50
Cold pdProj
13.87
14.87
14.87
Warm-started pdProj
7.23
8.23
8.23
Appendix
Table 21: Iterations, gradient calls, and function calls for the box-constrained QPs.
Problem
(s,t)
Method
Iter.
Grad. Calls
Func. Calls
Portfolio
(50,5)
IPOPT
14.68
16.68
15.68
Cold pdProj
8.99
9.99
9.99
Warm pdProj
3.48
4.48
4.48
Portfolio
(200,20)
IPOPT
21.06
23.06
22.06
Cold pdProj
12.07
13.18
13.18
Warm pdProj
9.36
10.57
10.57
Appendix
Table 22: Iterations, gradient calls, and function calls for the portfolio and SVM benchmarks.
Problem
Method
Iter.
Grad. Calls
Func. Calls
Quadrotor
IPOPT
9.20
11.20
10.20
Cold pdProj
9.74
10.74
10.74
Warm-started pdProj
8.00
9.00
9.00
Appendix
Table 23: Iterations, gradient calls, and function calls for the nonlinear quadrotor control problem.
Problem
Method
Converged
Avg. Runtime
Box QP, n=200
pdLIP
100%
4.02 ms
IPOPT
100%
85.27 ms
Constrained QP RHS, n=200
pdLIP
72.4%
20.11 ms
IPOPT
100%
315.54 ms
Constrained QP ALL, n=200
pdLIP
85.2%
19.93 ms
IPOPT
100%
344.84 ms
Appendix
Table 24: Direct approximate-solution results on the larger QP benchmarks.
Problem
Method
Converged
Avg. Runtime
Convex RHS
pdLIP
95.3%
5.23 ms
IPOPT
100%
4.29 ms
Convex ALL
pdLIP
91.7%
4.59 ms
IPOPT
100%
4.43 ms
Nonconvex RHS
pdLIP
92.5%
5.31 ms
IPOPT
100%
4.37 ms
Appendix
Table 25: Direct approximate-solution results on the n=50 constrained benchmarks.