When Known Physics Helps Neural PDE Models: Residual Constraints Out-Regularize Generic Priors for Nonlinear Dynamics
Authors: Zahra Farazpay, Aniruddha Bora
Organizations: Department of Physics Louisiana Tech University Ruston, LA, USA · Department of Computer Science Texas State University San Marcos, TX, USA
Neural PDE surrogates increasingly incorporate structural priors, yet it is often unclear whether their gains arise from physics-specific information or simply from regularization and training choices. We evaluate several such priors under a common protocol against a matched from-scratch neural operator baseline. Our central result is that a known-equation residual consistently outperforms the best generic regularizer at equal tuning budget. At fixed capacity this benefit appears across linear and nonlinear PDEs, but a capacity sweep reveals a sharp distinction: the advantage persists and grows for Burgers, KdV, and Allen-Cahn, while collapsing toward or below parity for linear heat and advection-diffusion. Thus, the durable value of the residual is specific to nonlinear operators. We further falsify a pre-registered hypothesis that the benefit is activated only by data sparsity: the residual remains advantageous even under full supervision. Its usefulness does, however, have a clear boundary. Under grid under-resolution, nonlinear coarse fields no longer satisfy the naive governing-equation residual, and enforcing it becomes actively harmful. In contrast, cross-family pretraining and in-context conditioning fail to outperform the strong from-scratch baseline in the regime studied. Together, these results identify when known physics provides non-redundant information to neural PDE models, when it does not, and when enforcing it introduces bias.
Figures & tables
Figure 1: Pushforward, not physics, strengthens rollout stability. One-step versus multi-step (pushforward) training across three viscosity regimes (5 seeds); pushforward reduces rollout error by 30–41%. This is a physics-free baseline-strengthening control; the headline residual and capacity comparisons remain matched one-step experiments.
Figure 2: Known-equation residuals provide a capacity-robust advantage for nonlinear operators. (a) Advantage Egeneric/Ephysics versus observation density ρ at fixed width; values above one favor the physics residual. The residual wins across all tested densities, including full supervision. (b) At full supervision, increasing width separates linear and nonlinear behavior: the nonlinear families retain a clear advantage while the linear controls approach or cross parity.
Figure 3: Illustrative Burgers rollout at sparse supervision ( ρ=0.08 ). The physics-penalized rollout preserves the shock more accurately than the best generic-regularized model (rollout error 5.2×10−3 versus 1.5×10−2 ). Quantitative multi-seed results are in Fig. 2 .
Figure 4: Non-redundancy predicts the durable physics-residual advantage (Prop. 1 ). Capacity-robust advantage Egeneric/Ephysics at the largest width versus the non-redundant content ρnr , the fraction of the one-step map no linear operator captures. The two linear families (open squares) sit at ρnr=0 and straddle parity (heat just above, advection–diffusion below); the three nonlinear families lie at ρnr>0 with durable advantages (Pearson 0.90 ). ρnr=0 coincides with the advantage collapsing to or below parity, identifying linearity rather than smoothness as the operative variable.
Figure 5: Resolution boundary. The true coarse-field residual stays at the discretization floor for linear families but grows rapidly under coarsening for nonlinear families, tracking unresolved tail energy. This is the failure mode of the naive physics penalty under under-resolution.
Figure 6: Cross-family transfer versus from-scratch training. In-context FiLM and attention-over-support remain nearly flat as target trajectories increase, while from-scratch FNO and U-Net models improve rapidly and overtake them by M=2 .
Figure 7: Transfer is more consistent with basin compatibility than source relevance across three held-out targets. For each held-out target (Burgers, KdV, Allen–Cahn), held-out rollout error vs. adaptation trajectories M for pretraining sources on a relevance gradient. Pretraining on relevant physics (green) fails to beat from-scratch (grey dashed) in the data-adequate regime and is a clear liability on all three targets; a rich non-physical operator (blue) is inert to marginally helpful; structureless pretraining (noise) is destructive. Related-PDE pretraining is worse than from-scratch in the data-adequate regime for all three targets.
Figure 8: 1D operators are resolution-invariant (scoped positive). Trained at N=128 , evaluated zero-shot at finer resolutions; rollout error is unchanged within ∼ 2% across five seeds.
Figure 9: On unseen geometry, locality beats the operator (scoped tension). Zero-shot rollout error on unseen obstacle layouts: a local CNN outperforms the FNO by ∼ 50% across five seeds with separated intervals. Locality and spectral structure trade off on different generalization axes.
Figure 10: Geometry, in the field. Truth, FNO, and CNN rollouts on an unseen obstacle layout (obstacles overlaid), with error fields and rollout error versus horizon. The local model resolves the wake; the operator smears it and its error grows over the rollout. Illustrative single case; the quantitative result is Fig. 9 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Direct verification of the closure bound and its refinement rate (Prop. 2 ). The dealiased Galerkin closure term τN=PN[F(PNu)−F(u)] computed on a fine reference grid. (a) ∥τN∥ versus the unresolved tail ∥QNu∥L2 : the nonlinear families fall on the slope- 1 line predicted by ∥τN∥≤C∥QNu∥∥u∥s (fitted slopes 0.99 , 0.98 , 1.01 in the bound’s norm), while the linear families collapse to a machine- ϵ cluster, τN≡0 . (b) For Burgers and KdV, ∥τN∥H−1 (solid) and ∥QNu∥ (dashed) decay in parallel under refinement, the closure rate matching the fitted spectral-regularity exponent σ .
Neural operators evaluate parametric partial differential equations cheaply but degrade sharply outside their training distribution. Physics-informed neural networks avoid dependence on labeled data, yet their optimization can be basin-fragile: when the governing residual admits multiple solutions, a PINN trained from scratch may converge to a physically incorrect state despite achieving a small residual. We show that these failure modes can be addressed jointly: an imperfect NO provides the structural prior needed to place a PINN in the correct solution basin, while the PDE residual refines the solution beyond the operator's accuracy. We introduce a three-stage framework that freezes the spatial basis of a physics-informed NO, extrapolates its solution branch to an out-of-distribution parameter using a polynomial continuation prior, and distills the resulting field into a fresh PINN. The NO need not be accurate at the target; it transfers solution-branch information, while PDE residual minimization in the PINN governs convergence. We evaluate the framework on three nonlinear PDEs: 1D viscous Burgers, 2D steady Allen-Cahn near a pitchfork bifurcation, and 2D steady lid-driven cavity flow. For Allen-Cahn, where the trivial solution satisfies the PDE residual exactly, a standard PINN collapses to the trivial zero branch, whereas distillation from the crude extrapolated operator recovers the non-trivial branch that matches the finite-difference reference. For the lid-driven cavity, extrapolating to a Reynolds number of Re = 3200 accelerates convergence to the correct physical state, achieving competitive accuracy using fewer parameters and optimization steps than recent literature baselines. These results establish a simple principle: an NO need not accurately predict the solution to be useful; it only needs to identify the correct basin from which PINN optimization can recover it.
S. Mohammad Mousavi, Teeratorn Kadeethum, Nikolaos Bouklas +1
Sibley School of Mechanical and Aerospace Engineering, Cornell University · AI Lab, Siemens Energy · Department of Civil and Systems Engineering, Johns Hopkins University
Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution. Recent hybrid methods promote physical correctness by targeting the PDE residual via gradient descent or Gauss--Newton steps, but inherit the compute cost and instability of the underlying classical optimizers. We show, theoretically and empirically, that numerically minimizing the PDE residual can be an unreliable proxy for reconstruction accuracy in ill-conditioned systems, explaining why these methods often do not make accurate predictions despite achieving low residuals. We propose error-conditioned Neural Solvers (ENS), built on a different principle: rather than an optimization target, the PDE residual field is passed as a direct input to the network at each iteration, enabling it to read the spatial structure of its own errors and learn an update policy to iteratively correct its predictions. Across four PDE families, ENS attains the highest prediction accuracy in the large majority of settings, with gains reaching 10× on turbulent Kolmogorov flow, while avoiding the expensive compute cost of hybrid methods. ENS's learned correction policy generalizes under distribution shift, including zero-shot parameter changes and cross-equation transfer, where its relative advantage is largest in the ill-conditioned regimes where residual minimization is least reliable. Project website: https://neuralsolver.github.io/.
Haina Jiang, Liam Wang, Peng-Chen Chen +4
University of Michigan · KAIST AI · Los Alamos National Laboratory
Neural operators are widely used as surrogate solution maps for partial differential equations (PDEs), but full-size models can be costly to store, deploy, and evaluate in many-query scientific workflows. This work introduces Operator Boosting, a stagewise residual-learning framework for constructing compact neural-operator surrogates directly, rather than training a large model and compressing it afterward. Starting from the empirical mean predictor in normalized output coordinates, the method trains a sequence of tiny same-family neural operators on residual fields and incorporates each correction through validation-selected shrinkage. We instantiate the framework with Fourier neural operators (FNOs), DeepONets, and convolutional neural operators (CNOs), and compare boosted tiny stacks against full-size monolithic baselines across one-, two-, and three-dimensional PDE benchmarks from PDEBench, APEBench, and The Well. Across 30 dataset-architecture pairs, 21 show positive mean accuracy gains and 17 have positive confidence intervals, while all boosted stacks reduce trainable parameter count by approximately 72-95%. Best-model comparisons show empirical Pareto improvements on 7 of 10 completed PDE benchmarks, including two-dimensional Navier-Stokes, shallow-water dynamics, Darcy flow, one-dimensional transport and reaction systems, and three-dimensional compressible Navier-Stokes. These results show that Operator Boosting often improves the empirical accuracy-parameter Pareto frontier of neural PDE surrogates, while also exposing PDE- and architecture-dependent regimes where residual boosting fails to offset compression.
Lennon J. Shikhman
Georgia Institute of Technology College of Computing Atlanta, Georgia, USA