When Known Physics Helps Neural PDE Models: Residual Constraints Out-Regularize Generic Priors for Nonlinear Dynamics
Organizations: Department of Physics Louisiana Tech University Ruston, LA, USA · Department of Computer Science Texas State University San Marcos, TX, USA
Abstract
Neural PDE surrogates increasingly incorporate structural priors, yet it is often unclear whether their gains arise from physics-specific information or simply from regularization and training choices. We evaluate several such priors under a common protocol against a matched from-scratch neural operator baseline. Our central result is that a known-equation residual consistently outperforms the best generic regularizer at equal tuning budget. At fixed capacity this benefit appears across linear and nonlinear PDEs, but a capacity sweep reveals a sharp distinction: the advantage persists and grows for Burgers, KdV, and Allen-Cahn, while collapsing toward or below parity for linear heat and advection-diffusion. Thus, the durable value of the residual is specific to nonlinear operators. We further falsify a pre-registered hypothesis that the benefit is activated only by data sparsity: the residual remains advantageous even under full supervision. Its usefulness does, however, have a clear boundary. Under grid under-resolution, nonlinear coarse fields no longer satisfy the naive governing-equation residual, and enforcing it becomes actively harmful. In contrast, cross-family pretraining and in-context conditioning fail to outperform the strong from-scratch baseline in the regime studied. Together, these results identify when known physics provides non-redundant information to neural PDE models, when it does not, and when enforcing it introduces bias.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| family | ratio | CI on | ||||
|---|---|---|---|---|---|---|
| heat (lin.) | 0.03 | 5.41 | [6.1e-07, 2.6e-06] | |||
| heat (lin.) | 0.06 | 4.96 | [5.2e-07, 2.4e-06] | |||
| heat (lin.) | 0.12 | 4.65 | [4.8e-07, 2.2e-06] | |||
| heat (lin.) | 0.25 | 4.69 | [4.5e-07, 2.2e-06] | |||
| heat (lin.) | 0.5 | 4.63 | [4.3e-07, 2.2e-06] | |||
| heat (lin.) | 1.0 | 4.58 | [4.3e-07, 2.2e-06] |
| Target | Metric | Related | Rich | No map | Noise | Scratch |
|---|---|---|---|---|---|---|
| KdV | Mean | 0.473 | 0.428 | 0.797 | 0.994 | 0.427 |
| 95% CI | [.464,.484] | [.412,.442] | [.758,.846] | [.941,1.05] | [.410,.443] | |
| Allen–Cahn | Mean | 0.185 | 0.048 | 1.184 | 1.370 | 0.055 |
| 95% CI | [.172,.198] | [.046,.050] | [1.04,1.32] | [1.20,1.54] | [.054,.056] |