PDE Operator Learning

PDE: Partial Differential Equation

Momentum

24 papers in the last four weeks, up 140% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 148

Oct 7, 2026cs.LG

CAFE+FNO: Fourier Kernel Generation via Multiplicative Feature Composition

The Fourier Neural Operator (FNO) learns solution operators of partial differential equations (PDEs) through Fourier-space kernel parameterization, but frequency truncation can limit the learning of high-frequency variations. AM-FNO and SirenFNO generate kernels for all grid modes from spectral coordinates using shared networks, making coordinate encoding and generator design important. Recent work on implicit neural representations (INRs) has proposed constructing frequency interactions through explicit feature composition rather than relying on subsequent MLPs to form them implicitly. Building on this approach, we propose CAFE+FNO, which incorporates Content-Aware Frequency Encoding+ (CAFE+) into Fourier kernel generation. CAFE+ combines Fourier--Chebyshev features through parallel affine branches and a Hadamard product, forming interactions within and across the two feature families. A kernel MLP maps the resulting representation of each normalized spectral coordinate to a complex channel-mixing matrix. Each layer shares its generator across all stored modes, making the number of trainable parameters independent of the number of modes for a fixed architecture. We compare CAFE+FNO with existing FNO variants on five PDE benchmarks and conduct ablation studies on basis configuration, multiplicative composition, and bandwidth learnability. Code and experimental configurations are available at https://github.com/fabsk101/CAFEPlusFNO.git.
Oct 6, 2026cs.LG

Learning PDE solution operators with variable initial conditions via Latent Dynamics Networks

In many-query scenarios, data-driven surrogate models provide an efficient alternative to high-fidelity solvers for simulating physical systems governed by Partial Differential Equations (PDEs). In this context, the Latent Dynamics Network (LDNet) has recently demonstrated remarkable performance in predicting the response of spatio-temporal systems, combining Neural Ordinary Differential Equations with nonlinear dimensionality reduction. However, the original formulation assumes a fixed initial condition, limiting its applicability to many real-world applications where a system evolves from varying starting states. In this work, we overcome this limitation while keeping the end-to-end training procedure of the original LDNet and its encoder-free nature, which preserves its intrinsic independence from spatial resolution and grid topology. We infer the initial latent state directly from a small set of early-time observations, treating latent-state initialization as an adaptation problem, and investigate two strategies: an auto-decoding formulation and a meta-learning approach in which the initial latent state acts as a task-specific context variable. We demonstrate the accuracy of the proposed methods across diverse physical phenomena, spanning advection-diffusion, fluid dynamics, and solid mechanics. Meta-learning markedly accelerates latent-state inference and induces smoother, better-conditioned optimization landscapes, and spontaneously organizes the latent space into a structured representation that reflects physically meaningful features of the underlying dynamics. The coordinate-based decoder enables training from spatially subsampled data while recovering high-resolution solution fields at inference. The resulting approach provides an efficient and resolution-independent surrogate modeling framework for many-query simulations of time-dependent PDEs with varying initial conditions.
Oct 5, 2026cs.LG

OCL-PDE: A Generative Framework for PDE Inverse Problems with Observation-Complementary Latents

Partial differential equation (PDE) inverse problems are often ill-posed, making fine-scale details difficult to recover. We address this problem by introducing a learned observation-complementary latent representation that preserves reconstruction-relevant information and is combined with the observation to reconstruct the unknown field. Building on this representation, we propose OCL-PDE, a generative framework that encourages the observation to guide large-scale structure and the latent to supply complementary fine-scale details. OCL-PDE is built on a physics-aware autoencoder (AE) and conditional Flow Matching, supporting inverse reconstruction as well as forward PDE prediction. Experiments demonstrate improved reconstruction accuracy and fine-detail recovery compared with the evaluated baselines.
Oct 5, 2026cs.LG

Transmission Factors for Lossy Compression of PDE Training Data: Measuring What Reaches a Trained Operator

Operator-learning benchmarks ship as full-precision arrays, and they have grown to terabyte scale. A curator who wants to distribute one has to decide how coarsely to store it. That decision is usually made by fixing a tolerance on the reconstruction error of the stored field. We show that this quantity is measured in the wrong place. It compares the stored field with the original, before any model is trained. What the curator is buying is the accuracy of a model trained on the compressed copy, and the two can disagree: two PDEBench families stored to identical field error differ threefold downstream. A solution operator smooths, so only part of the codec error ever reaches the model's output. That part can be measured. Push a compressed field through a model already trained at full precision, and compare its output against the output on the clean field. The measurement costs forward passes and no training on compressed data. It spans more than two orders of magnitude across six PDE families, and it orders downstream cost where field error does not, inverting 12 of 104 cross-dataset comparisons against field error's 36. Read as a budget, the same curve nominates a storage rate for each dataset. A back-check against trained models finds those rates wrong in both directions by up to a factor of two in stored bits.
Oct 4, 2026cs.LG

Poisson-GENERIC Neural Operators: Exact Metriplectic Structure in Function Space via Casimir Entropies

Existing thermodynamically consistent neural operators impose the GENERIC degeneracy conditions by projecting the reversible operator onto the complement of the entropy gradient. This makes the operator state-dependent and forfeits the Jacobi identity, so the result is metriplectic-degenerate rather than metriplectic. We instead obtain degeneracy the way GENERIC does. For nonlinear transport, the reversible operator LL is the compatible Lie-Poisson pencil αD+λ(uD+Du)αD+λ(uD+Du); otherwise it is a constant, trivially Poisson Fourier multiplier. On an augmented state (u,s)(u,s) with a latent entropy density, S=∫sS=\int s is a Casimir of LL, so L δS/δz=0L\,δS/δz=0 holds identically without projection. The energy combines a fixed mechanical quadratic, a learned gauge-free potential, and a convex internal energy. The friction operator M=AA⊤M=AA^\top satisfies M δE/δz=0M\,δE/δz=0 pointwise, and its Onsager parity structure permits diffusion and damping while provably excluding transport. For any parameters, skewness, positivity, both degeneracies, and the Jacobi identity (on the resolved band for the Lie-Poisson term) hold to machine precision. Heat conduction and damped waves admit exact closed-form friction operators, the second law bounds physical energy under a checkable curvature condition, and a discrete-gradient integrator yields exact discrete first and second laws. On four PDEs in 1D and 2D with three backbones (FNO, Transolver, CNO), the model wins 61 of 72 seed-level comparisons against same-backbone unconstrained baselines, learns the exact transport and wave symbols, matches the true dissipation rate within 13% on heat and Burgers, and dissipates nothing on advection. A constant-LL ablation isolates the cost of exact Jacobi as the loss of Burgers, while a learned-entropy ablation injects energy on every reversible-irreversible problem.
Oct 4, 2026cs.LG

Green-Routed Neural Operators:\Physics Determines Where the Network Reads

We identify a mismatch between the physical role of transport fields in many PDEs and their usual role in neural operators: PDEs use them to select read coordinates, whereas neural operators typically treat them only as input values. We address this mismatch with the Green-Routed Neural Operator (GRNO), which uses the governing equation to determine where latent features are sampled. A parameter-free equation adapter evaluates the diagnostic relation and constructs a departure map whose values are the read coordinates. A multiscale encoder-decoder combines centered and routed reads of latent features to learn the complete finite-time update. Across five two- and three-dimensional PDE systems, GRNO achieves the lowest mean final relative L2L^2 error on four under 40-step autoregressive evaluation and remains competitive on Keller-Segel. Fixed-weight route interventions reveal strong dependence on direction and spatial alignment in four systems, with weak dependence in Keller-Segel. In independently trained ablations, GRNO achieves lower mean errors than variants that supply the transport field only as an input feature, substitute a learned displacement for the equation-specified route, or apply the route with a spatial misalignment, across all five systems. It also substantially outperforms directly advecting the physical state and learning the remaining update, indicating that equation-specified read coordinates provide an effective structural prior for long-horizon PDE forecasting.
Oct 4, 2026cs.LG

METRO: Metric-Enhanced Token Routing Operator

State-of-the-art neural operators scale to complex meshes via slice-and-process architectures, yet many rely on linear compatibility scores for latent tokenization. Under common feature normalization, such scores are equivalent to isotropic Euclidean clustering, while without normalization they induce unbounded linear decision regions. In both cases, they lack slice-specific anisotropic locality, which can lead to redundant and entangled latent slices. To address this, we propose Metric-Enhanced Token Routing Operator (METRO), a geometry-aware routing mechanism that replaces linear projection with a learnable Mahalanobis metric. By enabling each latent slice to learn a local anisotropic tensor, METRO shapes receptive fields into exponentially localized, oriented ellipsoids that naturally align with flow features like boundary layers and wakes. As a drop-in replacement, METRO yields consistent improvements across both Transformer and Mamba backbones. Empirically, our method achieves substantial performance gains on irregular domains, outperforming baselines on both standard PDE benchmarks and complex industrial design tasks. Finally, METRO exhibits enhanced robustness in out-of-distribution regimes across varying Reynolds numbers and geometric configurations.
Oct 4, 2026cs.LG

Advectra: Asymmetric Latent Transport for Non-Stationary Physics

Many latent neural operators represent input and output fields in a stationary latent chart. In particular, common latent routing mechanisms use fixed or shared assignment weights for feature projection and reconstruction, limiting their ability to model transport-dominated systems where coherent structures move relative to fixed coordinate frames. We propose Advectra, a transport-aware latent operator that introduces a regularized kinematic coordinate map to decouple source and target coordinate systems. This yields an approximately co-moving latent reference frame and enables asymmetric feature aggregation and reconstruction. Combined with a geometry-aware ordering mechanism for state-space models, Advectra captures advective dynamics while maintaining stable global interactions. Advectra achieves the best performance among evaluated geometry-constrained and form-free baselines on advection-dominated benchmarks, including passive scalar transport in Navier--Stokes flows and Rayleigh--Taylor instability, while demonstrating strong generalization on real-world engineering tasks. These results highlight the benefit of explicit moving-frame structure in neural operators for non-stationary physics.
Oct 1, 2026cs.LG

Learning PDE Dynamics between Submanifolds Using Green's Observation Operators

Many physical systems are driven and observed only on lower-dimensional submanifolds of a larger spatial domain, while their dynamics are governed by the ambient medium occupying that domain. Examples include laser-heated parts imaged by an infrared camera, and ground-level emissions measured on a sensor plane. Full-domain solvers, however, compute the entire volume for every new source although only the observation submanifold is needed, and black-box surrogates do not exploit that the ambient medium remains fixed. We introduce the \emph{Green's Observation Operator (GObO)}, which maps the ambient medium once to the Green's kernel of a linear PDE restricted to the source and observation submanifolds. New sources then cost one lower-dimensional integral and no network evaluation. Exponential rates in the kernel yield an exact finite streaming state with horizon-independent memory; we prove its stability and an approximation rate for the restricted heat kernel. On three-dimensional heat conduction and advection--diffusion with collocated and distinct source and observation geometries, GObO trained on static sources predicts responses to moving sources zero-shot with 4--8×\times lower error than black-box surrogates, at 1.4,ms per query after a single conditioning pass. The same kernel transfers across resolutions and admits corrections for mild nonlinearities, including radiative losses and temperature-dependent conductivity, without retraining, at the cost of lower in-distribution accuracy.
Sep 30, 2026cs.LG

Beyond Simulation: Retain-and-Repair Neural Operators for Real-World Adaptation

Neural operators increasingly benefit from pretraining on numerical simulations, yet adapting them for real-world prediction remains challenging. We introduce the Retain-and-Repair Neural Operator (R2^2NO), a framework for adapting simulation-pretrained operators to real-world data while retaining useful pretrained structure. The pretrained operator is first finetuned on real data and then frozen to provide a source prediction, and a shared repair module learns a sequence of refinements from the same observations. Using orthogonal Fourier projections, a spectral ensemble fits a small ridge regression within each cell of the Fourier domain and combines the refinements by weights fitted on a held-out split of the real data. The cells are defined jointly by radial ranges, angular sectors, and measured channels, allowing refinement depth to vary with frequency magnitude, with orientation, and across channels. Including the source prediction as a candidate makes retention available in every cell, and independently trained repair modules enter the same combination as additional candidates. On all RealPDEBench systems and six backbones, R2^2NO consistently outperforms full finetuning and iterative refinement. The framework treats adaptation depth as a cell-specific choice learned from real data.
Sep 29, 2026cs.LG

Geometry-physics confounding impairs PDE learning across varying domains

Learning partial differential equation (PDE) dynamics across varying domains is central to predictive modelling and data-driven discovery of governing equations. However, geometric variation alters both field representation and the governing differential operators, confounding geometric effects with intrinsic physical properties in the observed dynamics. This work identifies geometry-physics confounding as a unified failure mechanism for PDE learning across varying domains. In forward operator learning, this confounding increases the burden of inferring geometry-dependent operator changes from finite data, reducing data efficiency and generalisation. In equation discovery, omitting geometry-induced operators misspecifies the candidate library, leading to biased parameters, missed governing terms and spurious terms. We propose a de-confounding framework that makes the known geometry-to-operator transformation explicit. Geometry-induced coefficient fields improve prediction and data efficiency across five operator-learning benchmarks, while geometry-complete candidate libraries recover the generating equations and reduce held-out PDE residuals by more than two orders of magnitude in both evolving-domain systems. By separating known geometric action from intrinsic physics, the proposed framework supports more reliable and data-efficient PDE learning across scientific and engineering problems with varying geometries.
Sep 29, 2026cs.AI

Transolver-σσ: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving

Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver based only on physical-state modeling can achieve lower one-step error, whereas its spectral-only counterpart can become more accurate at later rollout steps. Motivated by this observation, we present Transolver-σσ, a neural PDE solver based on joint spectral--physical subspace modeling. Within each block, adaptive physical-state interactions and spectral transformations are modeled in dedicated latent subspaces, whose responses are recomposed to enable information exchange between the two representations. Within the physical subspace, we introduce Slice-Residual Physics-Attention (SRPA), which preserves an explicit slice-space identity path while retaining learnable cross-slice interaction. In parallel, an axis-factorized Fourier operator captures global spectral structure. Across five well-established PDE benchmarks spanning steady-state prediction and time-dependent dynamics, Transolver-σσ achieves state-of-the-art with a benchmark-averaged relative error reduction of 33.4% over the strongest baseline for each metric, while consistently improving autoregressive rollout over single-operator counterparts. Transolver-σσ further delivers strong gains on coupled multiphysics systems and real-world fluid and combustion measurements from RealPDEBench, demonstrating its effectiveness beyond standard simulation benchmarks.
Sep 28, 2026cs.LG

Neural Harmonic Measure Operator

We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes the density of this measure as a transformer-based boundary kernel supervised by Walk-on-Spheres exit samples, so one trained kernel handles different boundary values on a shape with no retraining. We extend it to Poisson via a classical decomposition, with an auxiliary network amortizing the source-induced correction and avoiding the singular volume quadrature that breaks direct evaluation. At inference, new boundary values and new sources both yield PDE solutions by re-integration against the fitted kernel and lift, with no retraining. NHMO improves over four prior baselines on the MCB-B 3D variable-shape Poisson benchmark across all five categories, and is competitive with major neural-operator baselines on a controlled 2D testbed.
Sep 27, 2026cs.LG

Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning

Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the first time into a specifically designed variational autoencoder (VAE) structure, to solve a general class of nonlinear PDEs on smooth manifolds. AFDONet is the first neural PDE solver whose architectural and component design is fully guided by an established mathematical framework (in this case, AFD theory), turning neural operator design from an art to a science. Thus, AFDONet also exhibits exceptional mathematical explainability and groundness, and enjoys several desired properties. Furthermore, AFDONet achieves outstanding solution accuracy and competitive computational efficiency in several benchmark problems. In particular, thanks to its deep connections with AFD theory, AFDONet shows superior performance in solving PDEs on i) arbitrary (Riemannian) manifolds, and ii) datasets with sharp gradients. Overall, this work presents a new paradigm for designing explainable neural operator frameworks.
Sep 27, 2026cs.LG

Dynamic Kuramoto-Hodge Operators for PDEs on Complex Geometries and Topologies

Learning PDE operators on complex domains requires capturing interactions among fields on vertices, edges, and faces, alongside global responses shaped by topology. Existing neural operators accommodate irregular geometries but often overlook these distinct field supports or their condition-dependent coupling. We introduce the Dynamic Kuramoto--Hodge Operator (DKHO), which combines topology-constrained interactions with learned coordination. DKHO encodes conditions on their native cochain supports, evolves Kuramoto-inspired relation states through the boundary and coboundary operators that compose the Dirac operator, and decodes non-harmonic and harmonic responses in orthogonal Hodge subspaces. Topology thus determines where information can flow, while learned dynamics adapts how it is exchanged to each PDE instance. Across porous-medium Darcy flow, torus transport--diffusion, and cavity magnetostatics, DKHO-large reduces prediction error by approximately 61% on average over leading baselines, while DKHO-small remains competitive using only 11.5--24.3% as many parameters. These results suggest that coupling topological structure with adaptive dynamics provides an effective inductive bias for accurate and parameter-efficient PDE operator learning on complex geometries and topologies.
Sep 24, 2026math.NA

Elucidating the Conformal Structure of the Brinkman Penalisation Method for Geometry-Adapted, Structure-Preserving Operator Learning of Hamiltonian PDEs

The Brinkman penalisation method embeds boundary-value problems on complex domains into a simple computational box by modeling the solid region as a strongly dissipative medium, avoiding body-fitted mesh generation. We show that multi-symplectic Hamiltonian PDEs regularised by Brinkman-type penalisation retain a multi-conformal symplectic structure under a compatibility condition linking the symplectic matrix and the penalisation projection. This yields an exact local conservation law, under which the multi-symplectic two-form is conserved in the fluid region and decays exponentially inside the solid. The linear wave equation with Brinkman friction and Maxwell's equations with artificial Ohmic conductivity satisfy this condition, with explicit modified Hamiltonian densities. Building on this, we propose (i) structure-preserving numerical integrators via Strang splitting that satisfy a discrete conformal conservation law, and (ii) conformal symplectic neural operators that interleave exact dissipative flows with learnable multi-symplectic evolution operators, allowing geometry-dependent operator learning. Numerical experiments on wave and electromagnetic scattering demonstrate that our methods reproduce correct local energy budgets and avoid unphysical energy drift, providing a principled framework for physics-consistent scientific machine learning on complex domains.
Sep 24, 2026cs.LG

Physics and Data Driven Transformer-Mamba Framework for Flow Field

While deep learning accelerates expensive partial differential equation solving in computational fluid dynamics (CFD), existing methods like PINNs and FNOs often struggle with generalization, noise robustness, and physical consistency. We introduce the Transformer-Mamba for Flow Field (TM4FF) framework, a physics-constrained operator learning model with three key innovations: a Residual Wavelet Mamba (RWM) layer for feature denoising, a Transformer-based attention mechanism for enhanced feature fusion, and a physics-informed loss using Fourier derivatives to enforce the Navier-Stokes equations. Experiments on four CFD datasets show TM4FF achieves high accuracy and robust generalization across varying flow conditions.
Sep 23, 2026cs.LG

MENO: Memory-Efficient Neural Operator

We propose the Memory-Efficient Neural Operator (MENO) as a high-performance PDE neural solver based on the Manifold Function Encoder (MFE). MENO features three primary advantages: (1) MENO has a significantly smaller memory footprint and much faster training speed than other popular architectures, with the memory footprint being independent of the data resolution, and therefore holds the potential for scaling up to large-scale models. (2) MENO can accept PDE inputs of arbitrary form, including arbitrary geometric domains and arbitrary discretizations. In particular, it is capable of handling cross-geometry scenarios, i.e., where the input functions and the output solutions are defined on different manifolds. (3) MENO exhibits strong generalization capability, and achieves the best accuracy on most of the benchmarks we tested, compared with the results reported in the literature. The code is available on GitHub at https://github.com/jpzxshi/MENO, and all numerical examples in this paper can be run with a single command to reproduce the reported results.
Sep 20, 2026cs.LG

Predicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition

Neural operators have emerged as powerful surrogates for solving partial differential equations (PDEs), yet their reliability under distribution shift remains a critical barrier to deployment. Existing approaches to out-of-distribution (OOD) generalization in operator learning are largely empirical and black-box: they report aggregate error metrics without explaining why errors arise or when they will grow. We propose a structure-preserving framework that makes OOD generalization predictable and auditable. Our key idea is to parameterize the learned solution operator as a spectral filter hθ(λ)h_θ(λ) acting on the eigenvalues of the underlying elliptic operator, implemented via Chebyshev polynomial expansions and trained with a weak-form objective. This parameterization admits an exact decomposition of the energy-norm error into two observable components: a model-dependent spectral approximation term and a distribution-dependent spectral weighting term induced by the input. From this decomposition we derive three diagnostics: a conservative in-band supremum \varepssup\vareps_{\mathrm{sup}}, a global RMS proxy \varepsrms\vareps_{\mathrm{rms}}, and a sample-dependent effective metric \varepseff(f)\vareps_{\mathrm{eff}}(f). These diagnostics can be computed without access to ground-truth solutions. Through four controlled experiments, we show that \varepseff(f)∥f∥\vareps_{\mathrm{eff}}(f)\|f\| consistently predicts energy error under in-distribution, in-band spectral shift, out-of-band tail, and compound shifts, whereas global metrics can be systematically misleading. Our framework shifts OOD assessment of neural operators from black-box benchmarking to operator-structure diagnostics, providing a practical route to auditable scientific machine learning.
Sep 16, 2026math.NA

HiLNO: A Hierarchical Latent Neural Operator with Multi-Scale Supervision for PDEs on General Geometries

Latent neural operators improve the efficiency of operator learning for partial differential equations (PDEs) by performing the main computation on compact latent representations. However, directly compressing the input representation to obtain such compact representations may discard solution-relevant spatial information, especially for PDE solutions with multiscale structures. To address this problem, we propose HiLNO, a hierarchical latent neural operator that constructs a fine-to-coarse-to-fine latent space and further introduces multi-scale supervision (MSS) and anisotropic Gaussian attention. The hierarchy mitigates potential information loss during compression, while MSS aligns intermediate predictions with downsampled target fields, encouraging solution-relevant structures to be captured across multiple spatial scales. Anisotropic Gaussian attention enables feature transfer across the hierarchy, making HiLNO applicable to general geometries. Experiments on representative PDE benchmarks and a large-scale automotive aerodynamics task show that HiLNO achieves competitive predictive accuracy, while reducing the parameter count by an average of 84.4% and FLOPs by an average of 69.2% compared with LinearNO. Additional experiments demonstrate effective generalization to unseen spatial resolutions. Code is available at https://github.com/JcLimath/HiLNO.
Sep 15, 2026cs.LG

Lecture notes on Physics Informed Neural Networks, Neural Operators, and their applications

This is the set of lecture notes for the PhD course \href{https://www.unibz.it/en/faculties/engineering/phd-computer-science/study-course-offering/2025/36967}{\textit{Physics Informed Neural Network}, held at the University of Bozen/Bolzano} in the academic year 2025/2026. The goal of the course was to introduce the concept of Physics Informed Deep Neural Networks (PINN) and Neural Operators (NOs), discuss their implementation from scratch in PyTorch and using advanced ad-hoc developed open-source libraries such as NVIDia PhysicsNeMo to address real-world problems in various fields (engineering, physics, petroleum reservoir). We discuss recent topics such as Mixture-of-Models, Fourier Neural Operators, Physics-Informed Kolmogorov-Arnold Networks (PIKANs) and Fourier Neural Operators.
Sep 15, 2026cs.LG

Can Deep Learning Achieve Cross-Physics Mapping?

Can deep learning translate physical fields governed by fundamentally different equations? We address this question by introducing Cross-Physics Mapping (CPM), an operator-learning framework for mappings between heterogeneous physical domains. We formulate sufficient conditions for such mappings through compatible latent representations and propose a dimensionless scaling principle that aligns the characteristic evolution scales of the source and target systems without assuming their dynamical equivalence. As a representative test, paired diffusion and wave fields are generated independently from their respective parabolic and hyperbolic equations while sharing the same latent geometry, material heterogeneity, excitation, and dimensionless scale. Seven architectures-ResUNet, DeepONet, Fourier, latent, wavelet, U-shaped, and Galerkin neural operators-are evaluated for both diffusion-to-wave and wave-to-diffusion mappings. The results reveal a strong directional asymmetry. Diffusion-to-wave reconstruction is more challenging because it requires recovering wavefront, phase, and time-of-flight information attenuated by diffusion; U-NO performs best in this direction, achieving a relative ℓ2\ell_2 error of 0.3070.307 and an R2R^2 of 0.9050.905. Wave-to-diffusion mapping is considerably more stable, with GNO attaining a relative ℓ2\ell_2 error of 0.1540.154 and an R2R^2 of 0.9350.935. Neural operators generally outperform the conventional convolutional baseline, highlighting the nonlocal nature of cross-physics transformations. These findings demonstrate that deep learning can establish useful mappings between distinct physical modalities on a shared latent manifold, while the achievable accuracy remains fundamentally constrained by the direction-dependent information content of the governing physics.
Sep 15, 2026cs.LG

Stable by Construction: Variational Latent Markov Operators for Long-Horizon PDE Prediction

Neural PDE solvers provide efficient surrogates for time-dependent physical systems, but autoregressive prediction over long horizons remains challenging because local errors can induce distribution shift and accumulate under recursive deployment. We develop a variational approach to this problem by introducing latent Markov dynamics in which physical states are represented by latent distributions and evolved through probabilistic transitions. The framework is formulated directly on function spaces and specialized to functional Gaussian models, where structured latent perturbations induce a spectral geometry and variational transition alignment regularizes the learned dynamics. We further analyze how these mechanisms affect autoregressive error propagation, providing a theoretical connection between variational training and long-horizon prediction. We instantiate the framework as the Variational Autoencoding Markov Operator (VAMO), which combines spatially resolved latent fields, structured Gaussian perturbations, and a neural-operator transition. Empirically, we demonstrate the effectiveness of VAMO on several fluid-dynamics benchmarks with prediction horizons extending substantially beyond those represented during training, where it consistently reduces error accumulation and improves rollout stability over several deterministic and noise-injection baselines. Overall, these results highlight variational modeling as a complementary approach to robust long-horizon neural PDE dynamics.
Sep 14, 2026cs.LG

Single-condition neural solvers encode transferable response spaces for parametric differential equations

Operator learning for parametric partial differential equations (PDEs) typically builds global models over prescribed domains, requiring cross-condition data or costly physics-constrained training. Here we show that the output Jacobian of a neural solution model trained at one condition defines a reusable response space for cross-condition solution variations. We introduce Linearized Subspace Transfer (LST) to exploit this space and recover target solutions by minimizing the target PDE-system residual over response-space coordinates. Because any single response space has finite coverage, Active Transfer Modeling (ATM) uses post-transfer residuals as coverage indicators to selectively acquire response spaces from additional single-condition models. Across six systems, single-condition response spaces supported cross-condition transfer, with enrichment improving accuracy when added spaces expanded representation capacity. Relative to evaluated physics-informed operator baselines, ATM reduced error and offline construction cost, with orders-of-magnitude accuracy gains in representative cases and millisecond-to-second target adaptation. These results establish neural solvers as reusable local parametric models.
Sep 14, 2026cs.LG

Physics-Informed Conformal Prediction: Embedding PDE Consistency into Distribution-Free Uncertainty Quantification for Neural Operators

Neural operators such as the Fourier Neural Operator (FNO) achieve remarkable accuracy in approximating solutions to partial differential equations (PDEs). However, providing rigorous uncertainty estimates remains an open challenge. We propose Physics-Informed Conformal Prediction (PI-CP), a framework that embeds PDE residuals into the nonconformity score of split conformal prediction, producing prediction intervals that are (i) distribution-free with provable coverage guarantees, and (ii) spatially adaptive when the PDE residual correlates with prediction error -- tighter where physics is well-satisfied, wider where it is violated. Additionally, we prove that FNO's translation equivariance creates a fundamental approximation barrier for PDEs with Dirichlet boundary conditions, and show that coordinate channels resolve this with up to 63x error reduction. We validate PI-CP across six physics scenarios -- heat conduction (2D/3D), structural mechanics (2D/3D), Darcy flow, and Navier-Stokes -- demonstrating consistent 89-91% coverage for all four Conformal methods, while MC Dropout and Deep Ensembles are unstable (82-100%). FNO outperforms CNN and DeepONet by 10-12x.
Sep 14, 2026cs.AI

DU-NO: A Parameter-Efficient Double U-Shaped Neural Operator for Phase-Resolving Wave Modeling

Phase-resolving wave models such as FUNWAVE-TVD are the accuracy standard for nearshore dynamics, resolving the shoaling, refraction, and breaking of individual waves, but their cost rules them out for the ensembles, uncertainty quantification, and real-time warning that operational forecasting demands. Neural operators promise solver-level accuracy at a fraction of that cost, yet on wave-dominated fields the accurate ones are large: hybrid spectral-convolutional operators such as U-FNO (the strongest baseline in our study after DU-NO) buy their fidelity with tens of millions of parameters. We introduce DU-NO (Double U-shaped Neural Operator), a multiscale U-shaped spectral operator that attaches lightweight convolutional U-Net branches only at its two shallowest encoder and decoder levels. The placement follows a sampling argument: high-wavenumber content exists only on fine grids, so the local, full-band pathways go where that content lives, while the coarse, band-limited levels stay purely spectral. A depth-decaying mode schedule holds the model to 3.64M parameters, an order of magnitude below U-FNO. On our publicly released FUNWAVE-TVD benchmark, DU-NO attains the best autoregressive rollout error of six identically trained architectures, improving on U-FNO by 14.9% with 10.8x fewer parameters, and a frequency-band analysis shows the gain holds across all bands, including the high-wavenumber band where truncated-spectral operators collapse. Parameter-matched controls confirm the gain is architectural: rescaled to the same 3.6M budget, the best baseline still trails DU-NO by 28.6%. The advantage carries beyond nearshore waves: DU-NO matches the strongest baselines on 2D Navier-Stokes and wins clearly on PDEBench shallow-water rollouts. Code, trained models, and evaluation artifacts are available at https://anonymous.4open.science/r/duno-code-5A7B/.
Sep 7, 2026cs.LG

Two-Scale Localized PCA-Net: Coarse-Global and Local-Residual Representations for Artifact-Reduced PDE Operator Learning

Localized dimensionality reduction improves the scalability of operator learning for high-dimensional partial differential equations (PDEs), but independently decoded local patches can introduce block offsets, interface mismatches, and spurious high-wavenumber content. We introduce Two-Scale Localized PCA-Net, which decomposes the solution into a coarse-global component and local residual corrections. A compact global PCA basis captures domain-scale structure, while nonoverlapping local PCA bases represent the remaining fine-scale residual. A block-balanced latent objective couples the two representations, and optional interface-aware fine-tuning further promotes continuity through reconstruction and trace losses. On Poisson benchmarks, the two-scale representation substantially reduces reconstruction error and visible block artifacts relative to plain and overlap-based localized PCA-Net while approximately halving PCA fitting cost relative to overlap. On heterogeneous Darcy flow, it strongly reduces interface and discrete-residual errors, with more modest reconstruction gains. Ablations show that the primary improvement arises from the two-scale output representation, while interface-aware fine-tuning provides complementary continuity refinement. Overall, separating globally coherent structure from localized residual detail provides an efficient representation for artifact-reduced PDE operator learning.
Sep 7, 2026cs.LG

Local gradient neural operator

Field temporal prediction and source identification constitute canonical problems in dynamical systems. Conventional approaches to these problems depend on a thorough understanding of the governing partial differential equations (PDEs). Recently, deep learning, as represented by neural operators, has provided a data-driven paradigm for addressing such tasks. However, most existing global neural operators for PDEs require large training datasets and many learnable parameters, with limited interpretability and generalization. We propose the local gradient neural operator (LGNO) as a lightweight and interpretable alternative for field temporal evolution prediction and source identification in typical mechanical problems. The method builds on priors from nonlinear gradient discretization and uses multilayer perceptron convolutional layers to learn translation-invariant local kernels that resemble discrete stencils. A zero consistent stencil factorization separates coefficient learning from field reconstruction, rendering the learned operators more transparent. For problems with symmetries, network folding shares equivalent components and reduces parameter counts. We evaluate the method on PDE benchmarks covering linear and nonlinear, static and dynamic, and low and high dimensional cases. Results show that LGNO maintains accuracy, parameter efficiency, and rollout stability across these tasks, and further exhibits wide applicability to mechanical problems including diffusion, flow, and quantum phenomena.
Sep 2, 2026cs.LG

Equation Recast for Canonical Operator Learning Across Parametric PDEs

Learning solution operators across broad parameter ranges can require substantial coverage of both input functions and physical parameters, particularly for purely data-driven parametric models. In addition, the resulting models may fail silently outside the training distribution. We introduce equation recast, which reformulates parametric operator learning as the learning of a single canonical operator. Parameter-induced operator variations are derived analytically from the governing equation and absorbed into effective sources, enabling zero-shot prediction across new parameter regimes. Across multi-parameter, nonlinear, and singular PDE settings, equation recast supports extrapolation, integrates sparse heterogeneous datasets in a shared canonical representation, and uses loss of convergence as an internal warning signal for failure of the recast iteration. In high-fidelity tokamak simulations for nuclear fusion, the framework unifies electron-temperature data across four device geometries through canonical-domain mapping within one jointly trained operator. Equation recast provides a route toward reusable neural PDE solvers combining equation-guided transfer, data efficiency, and monitorable inference.
Sep 1, 2026cs.LG

VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows

Accurate prediction of unsteady separated flows is challenging because the aerodynamic loads depend on nonlinear separation and vortex-shedding dynamics. Although high-fidelity CFD resolves these mechanisms, its cost limits repeated use in design and control. Standard field-level surrogate training, however, does not distinguish the flow regions that contribute most strongly to the aerodynamic loads. We introduce VATO (Vortex-Force-Aware Transformer Operator), which couples the Vortex Force Map (VFM) method to a geometry-aware neural operator through two complementary mechanisms. VATO-S adds training-only supervision of the local VFM force-contribution field, with no increase in model size or inference cost. VATO-A uses VFM contribution and sensitivity fields to prioritise force-relevant source locations for residual cross attention. The methods are evaluated on unsteady CFD data for double-edged-plate aerofoils over 54 trajectories from nine geometries. Over lead times of 1-20~ms, VATO-S reduces velocity, pressure, and vorticity errors by 10.4%, 1.0%, and 15.6%, respectively, while VATO-A achieves reductions of 15.8%, 7.5%, and 31.2%. VATO-S gives the lowest VFM-derived drag error, whereas VATO-A gives the lowest pressure-derived lift and drag errors. Over lead times extending 50% beyond the training range, VATO-A retains a 26.9% reduction in vorticity error and larger improvements in all four force readouts, despite reduced gains in velocity and pressure. These results show that force-aware operator learning can improve both flow-field prediction and aerodynamic functional accuracy in unsteady separated flows.