Singular parameters and missing limits in neural PDE solvers
Authors: Daniel Fernández
Organizations: Chair for Dynamics, Control, Machine Learning, and Numerics (Alexander von Humboldt Professorship), Department of Mathematics, Friedrich-Alexander-Universität Erlangen-Nürnberg, 91058 Erlangen, Germany.
Neural solvers for partial differential equations (PDEs) can approach an accurate solution while their parameters grow without bound. In such cases, the limiting solution may have no finite representation in the chosen model, leaving the best loss unattained. Our analysis connects missing limits in deep neural tanh- networks to unbounded hidden parameters or increasingly redundant neurons. For a class of models built from translated kernels, we describe the missing functions and recover them by adding kernel derivatives to the model. This completion makes the best approximation attainable under standard assumptions. Numerical studies follow the associated parameter growth and explore how completion affects PDE optimization.
Figures & tables
Figure 1 . Two merging tanh neurons approach a function outside the model. (a) Weights grow with opposite signs; (b) large contributions nearly cancel; (c) their sum approaches u∗ .
Figure 2 . Network outputs can converge while parameters escape or neurons degenerate. The limiting state lies outside the model. Schematic.
Figure 3 . As three Gaussian centres merge, weights grow and the fit improves. A derivative block represents the target exactly. Top: individual contributions; bottom: their sum and the target.
Figure 4 . Two tanh neurons merge during training as their weights grow with opposite signs. Circles mark starts; triangles mark endpoints. (a) Single precision; (b) double precision.
Figure 5 . Large neuron contributions nearly cancel, bringing their sum closer to the target. Top: contributions; bottom: sum (dashed) and target (green). The last two columns show independent refinements.
Figure 6 . Larger coefficient budgets improve the Gaussian approximation. (a) States; (b) errors; (c) rescaled excess energies approach their predicted limits. Dotted line: completed error for ε=0.01 .
Figure 7 . Completion sharply reduces the Poisson error in a representative run. Top: target and final approximations; bottom: error maps. Larger circles indicate larger coefficients.
Standard optimization
Completion
Relative H1 error
9.8×10−6–6.2×10−5
1.7×10−8–2.6×10−7
median
1.8×10−5
8.1×10−8
Coefficient norm
2.3×102–5.0×103
3.53694–3.53698
Smallest center half-gap
2.1×10−4–3.1×10−3
0
Solves (median)
758–1785 (1204.5)
377–525 (431)
Completed pairs
0
4
Table 1 . Final accuracy, coefficient sizes, and work for the Poisson–Neumann problem over ten shared starts.
Figure 8 . Completion improves Poisson accuracy and controls coefficient growth across ten starts. Left: best relative error so far; right: coefficient norm. Crosses mark abnormal termination.
Standard optimization
Completion
Relative H1 error
2.1×10−6–3.5×10−5
3.4×10−8–1.8×10−5
median
1.7×10−5
2.3×10−7
Relative eQ
4.8×10−6–7.4×10−5
2.5×10−8–3.4×10−5
Coefficient norm
7.0×101–1.0×103
2.5×100–1.0×102
Smallest center half-gap
7.5×10−4–1.3×10−2
0
Fits (median)
149–681 (225.5)
108–713 (255.5)
Table 2 . Final accuracy, coefficient sizes, and work for the Matérn problem over ten shared starts.
Figure 9 . Completion improves Matérn accuracy and reduces coefficient growth. Left: best relative error so far; right: coefficient norm. Crosses mark abnormal termination.
Standard optimization
Completion
Problem
τ
Reached
Median
Reached
Median
Poisson
1.0×10−3
10/10
51.5
10/10
69
1.0×10−5
1/10
1283
10/10
291.5
1.0×10−7
0/10
—
7/10
427
Matérn
1.0×10−3
10/10
20
10/10
20
1.0×10−5
3/10
169
9/10
180
Table 3 . Work needed to reach each relative H1 error threshold. Counts are out of ten starts; medians include only starts that reach the threshold.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Standard optimization
Completion
Problem
Order
Median H1
≤10−7
Median H1
≤10−7
Poisson
32
1.8×10−5
0/10
8.1×10−8
7/10
64
1.9×10−5
0/10
7.9×10−8
6/10
Matérn
64
1.7×10−5
0/10
2.3×10−7
5/10
96
1.8×10−5
0/10
1.5×10−8
6/10
Appendix
Table 4. Completion retains lower median errors with finer training quadrature.
Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution. Recent hybrid methods promote physical correctness by targeting the PDE residual via gradient descent or Gauss--Newton steps, but inherit the compute cost and instability of the underlying classical optimizers. We show, theoretically and empirically, that numerically minimizing the PDE residual can be an unreliable proxy for reconstruction accuracy in ill-conditioned systems, explaining why these methods often do not make accurate predictions despite achieving low residuals. We propose error-conditioned Neural Solvers (ENS), built on a different principle: rather than an optimization target, the PDE residual field is passed as a direct input to the network at each iteration, enabling it to read the spatial structure of its own errors and learn an update policy to iteratively correct its predictions. Across four PDE families, ENS attains the highest prediction accuracy in the large majority of settings, with gains reaching 10× on turbulent Kolmogorov flow, while avoiding the expensive compute cost of hybrid methods. ENS's learned correction policy generalizes under distribution shift, including zero-shot parameter changes and cross-equation transfer, where its relative advantage is largest in the ill-conditioned regimes where residual minimization is least reliable. Project website: https://neuralsolver.github.io/.
Haina Jiang, Liam Wang, Peng-Chen Chen +4
University of Michigan · KAIST AI · Los Alamos National Laboratory
Uncertainty quantification for partial differential equations is traditionally grounded in discretization theory, where solution error is controlled via mesh/grid refinement. Physics-informed neural networks fundamentally depart from this paradigm: they approximate solutions by minimizing residual losses at collocation points, introducing new sources of error arising from optimization, sampling, representation, and overfitting. As a result, the generalization error in the solution space remains an open problem. Our main theoretical contribution establishes generalization bounds that connect residual control to solution-space error. We prove that when neural approximations lie in a compact subset of the solution space, vanishing residual error guarantees convergence to the true solution. We derive deterministic and probabilistic convergence results and provide certified generalization bounds translating residual, boundary, and initial errors into explicit solution error guarantees.
Amartya Mukherjee, Maxwell Fitzsimmons, David C. Del Rey Fernández +1
Department of Applied Mathematics, University of Waterloo, Waterloo, ON, Canada
Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain. We introduce \EMNGDfull{}, a manifold optimization framework for physics-informed and variational neural PDE solvers whose parameters lie on a Riemannian manifold. EMNGD restricts the energy-induced quadratic model to feasible tangent directions and uses retractions to preserve parameter constraints throughout optimization. Under coercivity, we prove that the push-forward of the undamped EMNGD direction is the best feasible approximation to the function-space Newton vector in the energy metric. We establish coordinate invariance, exact reduction to ENGD in Euclidean space, global first-order convergence with Armijo backtracking, and robustness to inexact tangent solves. For quadratic residual energies and generalized Gauss--Newton pullbacks, the Woodbury identity transfers the tangent system to sample space without changing the direction. Nyström approximation provides scalable sample-space solves with controlled direction error and recovers the exact direction after iterative convergence. On the evaluated neural PDE benchmarks, EMNGD achieves higher accuracy and faster convergence than the compared state-of-the-art baselines. Woodbury preserves the EMNGD direction, while scalable-solver diagnostics quantify the accuracy--cost trade-off of preconditioning and residual subsampling.
Zhangyong Liang, Huanhuan Gao
National Center for Applied Mathematics Tianjin University Tianjin, 300072, China · School of Mechanical and Aerospace Engineering Jilin University Changchun, 130025, China