cs.LGSep 27, 2026

The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning

Authors: Ameya D. Jagtap

Organizations: Aerospace Engineering Department, Worcester Polytechnic Institute, Worcester, MA 01609, USA

Abstract

Automatic differentiation (AD) lets neural networks compute derivatives of governing equations to machine precision, and this precision has made it the computational backbone of physics-informed machine learning. Yet exactness in the mathematical sense is not the same as fidelity to the physics. Here I argue that a derivative can be numerically perfect and still be the wrong derivative for the problem at hand, because AD, by construction, has no notion of the physical structure a solution must obey. Convection and its associated directionality, diffusion, and dispersion are only the most visible instances of a much longer list that spans all branches of computational science and engineering, including conservation, thermodynamic consistency, symmetry, symplectic structure, positivity, monotonicity, and boundedness. Recognizing this broader gap reframes how the field should build the next generation of PDE-driven neural surrogates.

Figures & tables

Explore similar work

Aug 11, 2026cs.LG

Derivative Computation in PINNs: Automatic Differentiation, Finite Differences and Beyond

We systematically investigate finite-difference (FD) derivative computation in Physics-Informed Neural Networks (PINNs) as an alternative to automatic differentiation (AD). On three benchmark PDEs we show that, with a properly calibrated step size, FD matches AD in accuracy on every problem while running faster across the full tested batch-size range and using substantially less GPU memory, and that a stochastic variant we propose outperforms AD on a stationary problem. We further show that for neural architectures with inter-sample dependencies (e.g. BatchNorm, self-attention) the standard PyTorch autograd idiom is silently incorrect; the correct per-sample alternative is computationally infeasible at PINN-relevant batch sizes, while FD provides a forward-only approximation that is empirically an order of magnitude closer to the true per-sample derivative.
Jun 13, 2026cs.LG

Automatic Differentiation from Scratch: How PyTorch Computes Gradients in Physics-Informed Neural Networks

This paper traces, with explicit numerical values, how PyTorch's automatic differentiation (AD) engine computes gradients for Physics-Informed Neural Network (PINN) training -- a setting that requires two levels of differentiation: computing the physics derivative y^′(t)=dy^/dt\hat{y}'(t)=d\hat{y}/dt through the network, and computing parameter gradients ∇θL\nabla_θL of a loss that itself depends on y^′(t)\hat{y}'(t). Using a 1-3-3-1 multilayer perceptron and the initial value problem y′(t)+y(t)=0y'(t)+y(t)=0, y(0)=1y(0)=1, we trace the complete pipeline at every node: the computational graph built during the forward pass, the reverse-mode backward traversal that computes all 22 parameter gradients in a single pass, and the graph-on-graph mechanism by which \texttt{create_graph=True} enables correct differentiation through the physics-informed residual. Every adjoint value is verified against the hand derivations of Tahimi (2026), connecting the P/QP/Q sensitivity framework to the vector--Jacobian products used by PyTorch's autograd engine.
Sep 7, 2026math.NA

A Systematic Analysis of Automatic Differentiation versus Discretization-based Constraints for Physics-Informed PDE Solvers

Physics-informed neural networks (PINNs) represent a growing frontier in using artificial intelligence to solve partial differential equations (PDEs). Automatic differentiation (AD) plays a central role in this paradigm, which is mesh-free and replaces traditional iterative solvers with gradient-based optimization in continuous space. However, the inherent limitations of AD, particularly in handling higher-order derivatives and discontinuous solutions, pose significant challenges for complex problems. This has motivated a growing number of researchers to explore discretization-based constraints as an alternative path. Yet, the respective applicability of these two paradigms remains largely unexplored. In this work, we conduct systematic experiments across a wide spectrum of problems, from simple linear Poisson to high-Mach hypersonic flows with strong discontinuities. Through a rigorous decomposition of approximation, optimization, and truncation errors, we systematically elucidate the fundamental trade-offs and error-governing mechanisms of both paradigms, as well as two representative network architectures: multi-layer perceptron (MLP) and graph neural network (GNN). Our results reveal a consistent trend: as nonlinearity strengthens, the accuracy advantage of discretization-based constraints becomes increasingly pronounced, with smaller optimization errors compensating for the truncation errors. Moreover, the more complex the nonlinearity and boundary conditions, the greater the advantage of GNN over MLP. These insights offer a robust practical guideline for configuring neural PDE solvers in demanding engineering applications. Our source data and code are available at https://github.com/guoxing0809/neuropde_analysis.