lapanda: A Matrix-Free Differentiable Solver for Nonconvex Constrained Optimization Layers
Authors: Yuankun Chen, Zifei Nie, Kangyu Lin, Ján Drgoňa, Liang Wu
Organizations: School of Artificial Intelligence, Jilin University, China · Graduate School of Informatics, Kyoto University, JPN · Department of Civil and Systems Engineering, Johns Hopkins University, USA
Differentiable optimization brings the structural guarantees of mathematical optimization to network pipelines, allowing them to be trained end-to-end. However, its application remains challenging for nonconvex constrained problems, as existing differentiable solvers often suffer from limited modeling expressiveness due to their reliance on specialized problem structures, while also incurring substantial computation time and memory overhead in both the forward and backward passes. To address these challenges, we propose lapanda, a matrix-free differentiable solver for nonconvex optimization with general constraints. It reformulates the problem to a sequence of augmented Lagrangian subproblems, each handled by a first-order inner solver through a proximal averaged quasi-Newton algorithm with adaptive linesearch, thus enabling efficient forward optimization. We establish local well-posedness of the solution map and convergence of the outer iterations, and further derive a sensitivity alignment between the original problem and the final subproblem in the backward pass, demonstrating that the subproblem sensitivity, which can be computed efficiently in a matrix-free manner, provides a principled approximation to the exact optimizer sensitivity. We evaluate lapanda on nonconvex constrained Rosenbrock benchmarks, imitation learning with several representative constrained optimal control problems, and embedded robotic obstacle-avoidance tasks. Compared with state-of-the-art differentiable solvers, lapanda delivers substantial reductions in computation time and memory footprint while maintaining reliable constraint satisfaction and learning performance.
Figures & tables
Solver
Target problem
Hard constraints
Matrix-free
Platform
lapanda (Ours)
General NLP
General
✓
C / MATLAB / Python
CasADi (2019)
General NLP
General
✗
C / MATLAB / Python
acados (2025)
General OCP
General
✗
C / MATLAB / Python
TurboMPC (2026)
General OCP
General
✗
Python
SafePDP (2021)
General OCP
General
✗
Python
PANDA (2026)
Composite NLP
Prox.-friendly
✓
MATLAB
Table 1: Comparison of representative differentiable optimization solvers.
Figure 1: Overview of the lapanda framework.
n
Forward time (ms)
Backward time (ms)
Memory (MB)
lapanda
Explicit
CasADi
lapanda
Explicit
CasADi
lapanda
Explicit
CasADi
100
8.23
–
36.50
4.04
4.22
3.04
1.78
5.01
17.55
200
14.37
–
108.19
5.89
10.23
17.05
2.06
7.78
29.39
500
49.75
–
496.03
14.36
51.26
245.99
3.12
25.32
84.70
1000
200.00
–
1968.78
37.36
218.21
2325.16
5.31
79.65
249.66
Table 2: Matched-accuracy comparison on the constrained Rosenbrock benchmark.
Figure 2: Normalized constraint values xi2+xi+12/ri(θ) for n=200 .
Figure 3: Open-loop imitation learning results on different constrained OCPs.
Figure 4: Closed-loop imitation learning results including training loss and policy evolution.
OCP
Method
Fwd. time (ms)
Bwd. time (ms)
Total time (ms)
Memory (MB)
CartPole
lapanda
5.10
0.09
5.20
1.64
SafePDP
22.06
9.65
31.71
7.71
Quadrotor
lapanda
1.20
0.24
1.45
1.79
SafePDP
29.86
13.41
43.27
9.23
TurboMPC-CPU
10.46
7.74
18.20
241.30
TurboMPC-GPU
21.78
2.28
24.06
336.48
Table 3: Detailed performance comparison of different solvers on closed-loop OCP benchmarks.
Optimization layers enable the incorporation of structured constraints and decision problems into learning systems. Training such systems requires differentiating through the embedded optimization problem, which can be challenging for general conic programs. We introduce dOPT, a solver-agnostic framework that, rather than differentiating the full conic formulation, reduces it at a computed primal-dual solution to an equality-constrained quadratic program that preserves the reference solution and its first-order sensitivity. The reduction captures the local first- and second-order conic geometry relevant to differentiation and remains well defined at singular configurations. Computing solution derivatives then requires a single symmetric linear solve, independently of the forward solver. We derive explicit reductions for convex NLPs, QPs, SOCPs, and SDPs. Numerical experiments validate the computed gradients and show favorable backward-pass scalability, with substantial speedups over existing differentiable conic optimization methods as problem size increases.
Fengyu Yang, Connor W. Magoon, Tyler Watts +1
Department of Mathematics University of North Carolina at Chapel Hill
Nonlinear Parametric Optimization Network (NLPOpt-Net) is an unsupervised learning architecture to solve constrained nonlinear programs (NLP). Given the structure of an NLP, it learns the parametric solution maps with guaranteed constraint satisfaction. The architecture consists of a backbone neural network (NN) followed by a multilayer (k-layered) projection. While the NN drives toward optimality through a loss function consisting of a modified Lagrangian augmented with a consistency loss, the projection ensures feasibility by projecting the NN predictions in the original constraint manifold. Instead of typical distance minimization, our projection exploits local quadratic approximations of the original NLP. Under certain conditions (such as convexity), the projection has a descent property, which improves the NN predictions further. NLPOpt-Net deploys an inversion-free, modified Chambolle-Pock algorithm to solve the constrained quadratic projections during the forward pass and uses the implicit function theorem for efficient backpropagation. The fixed structure of the projection further allows decoupling of the NN and the projection once the training is complete. NLPOpt-Net solves large-scale convex QP, QCQP, NLP, and nonconvex problems with near zero optimality gap and constraint violations reduced to machine precision. Additionally, it provides near accurate prediction of the active sets and corresponding dual variables, thereby enabling a scalable approach for multiparametric programming. Compiling the projection in C provides order of magnitude improvement in inference time compared to JAX. We provide the codes and NLPOpt-Net as a ready to use package that includes GPU support.
Deep unfolding (DU) accelerates iterative optimizers by introducing learnable components and training them through unrolled iterations, but extending DU to the large-scale semidefinite programs (SDPs) common in robotics has remained limited. Unrolling a full-update conic solver such as COSMO exposes two obstacles that prior work on learned conic solvers has not: backpropagating through the per-iteration linear-system solve incurs memory quadratic in the problem size once the coefficient matrix is formed explicitly, and backpropagating through the positive semidefinite (PSD) cone projection becomes numerically unstable when eigenvalues coincide. We address the first obstacle with a matrix-free implicit differentiation rule that operates entirely through matrix-vector products, reducing memory from O(n2) to O(n) and enabling backpropagation at scales where direct factorization runs out of memory. We address the second with a backward rule based on the Dalečkii--Krein representation of the Fréchet derivative, which remains well-defined under repeated eigenvalues. Together these make it possible to learn lightweight hyperparameter policies and warm-starts for a full-update conic solver. We evaluate on nonlinear covariance steering problems solved via sequential convex programming (SCP), as well as standalone SDPs and second-order cone programs ranging from max-cut and Lovász ϑ SDPs to robust estimation and control problems. The learned policies outperform state-of-the-art solvers across all problems, and can provide up to a 50× speedup depending on the class. When used as a subroutine in SCP, the learned approach delivers over a 30× speedup compared to COSMO.
Alex Oshin, Rahul Vodeb Ghosh, Evangelos A. Theodorou