The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning
Organizations: Aerospace Engineering Department, Worcester Polytechnic Institute, Worcester, MA 01609, USA
Abstract
Automatic differentiation (AD) lets neural networks compute derivatives of governing equations to machine precision, and this precision has made it the computational backbone of physics-informed machine learning. Yet exactness in the mathematical sense is not the same as fidelity to the physics. Here I argue that a derivative can be numerically perfect and still be the wrong derivative for the problem at hand, because AD, by construction, has no notion of the physical structure a solution must obey. Convection and its associated directionality, diffusion, and dispersion are only the most visible instances of a much longer list that spans all branches of computational science and engineering, including conservation, thermodynamic consistency, symmetry, symplectic structure, positivity, monotonicity, and boundedness. Recognizing this broader gap reframes how the field should build the next generation of PDE-driven neural surrogates.
Figures & tables
| Derivative type | Example | Can AD evaluate it? |
|---|---|---|
| Local, integer-order | Yes — exactly, via the chain rule | |
| Nonlocal, fractional or integral-order | No — requires evaluating an integral over the domain |
| Property | Representative strategies | Main trade-off |
|---|---|---|
| Directionality | Artificial diffusion [ 58 ] ; curriculum learning [ 19 ] ; time-domain decomposition [ 18 ] ; scalar or self-adaptive loss reweighting [ 59 , 60 ] ; hybrid numerical–AD stencils [ 62 , 63 ] ; hybrid solver [ 64 , 65 ] | Each restores an upwind bias or causal ordering by hand-tuned terms, schedules, or an explicit mesh, reintroducing the discretization choices AD was meant to avoid. |
| Conservation / thermodynamic consistency | Entropy-stable / kinetic-energy-preserving fluxes built into the loss or architecture [ 26 , 27 ] ; structure-preserving PINNs enforcing energy dissipation [ 67 ] | Requires a known discrete flux or free-energy form; demonstrated mainly on phase-field-type equations so far. |
| Invariance | Symmetry-loss penalties and invariant-surface conditions [ 68 , 69 , 70 , 71 ] ; hard-constrained or equivariant architectures [ 72 , 73 ] ; symmetry-reduced ansatz and operators [ 74 ] | Loss-based methods enforce symmetry only at sampled points; architectural methods are limited to symmetries with a known algebraic representation. |
| Symplecticity | Symplectic-by-construction networks [ 76 ] and pseudo-symplectic relaxations [ 77 ] ; hybrid PINN–symplectic integrators [ 78 ] ; symplectic neural operators [ 79 ] ; non-canonical / nearly-periodic structure-preserving networks [ 80 ] | Restricts the hypothesis class to (near-)symplectic maps, or shifts the structure-preservation burden onto a classical integrator whose accuracy depends on the learned Hamiltonian. |
| Positivity / monotonicity / boundedness | Hard-constrained output layers [ 81 , 82 ] ; structure-preserving PINNs for bounded order parameters [ 67 ] ; constrained-optimization / adversarial reformulation [ 83 ] | Requires the constraint form to be known analytically in advance, or introduces the optimization instability of adversarial training. |