Many latent neural operators represent input and output fields in a stationary latent chart. In particular, common latent routing mechanisms use fixed or shared assignment weights for feature projection and reconstruction, limiting their ability to model transport-dominated systems where coherent structures move relative to fixed coordinate frames. We propose Advectra, a transport-aware latent operator that introduces a regularized kinematic coordinate map to decouple source and target coordinate systems. This yields an approximately co-moving latent reference frame and enables asymmetric feature aggregation and reconstruction. Combined with a geometry-aware ordering mechanism for state-space models, Advectra captures advective dynamics while maintaining stable global interactions. Advectra achieves the best performance among evaluated geometry-constrained and form-free baselines on advection-dominated benchmarks, including passive scalar transport in Navier--Stokes flows and Rayleigh--Taylor instability, while demonstrating strong generalization on real-world engineering tasks. These results highlight the benefit of explicit moving-frame structure in neural operators for non-stationary physics.
Figures & tables
Figure 1 : Overview of the Advectra operator.
Model
AirfRANS
ShapeNet-Car
Volume ↓
Surf ↓
CL↓
ρL↑
Volume ↓
Surf ↓
CD↓
ρD↑
GraphSAGE [ 24 ]
0.0087
0.0184
0.1476
0.9964
0.0461
0.1050
0.0270
0.9695
PointNet [ 25 ]
0.0253
0.0996
0.1973
0.9919
0.0494
0.1104
0.0298
0.9583
Graph U-Net [ 26 ]
0.0076
0.0144
0.1677
0.9949
0.0471
0.1102
0.0226
0.9725
MeshGraphNet [ 6 ]
0.0214
0.0387
0.2252
0.9945
0.0354
0.0781
0.0168
0.9840
GNO [ 8 ]
0.0269
0.0405
0.2016
0.9938
0.0383
0.0815
0.0172
0.9834
Table 2 : Quantitative results on industrial benchmarks AirfRANS and ShapeNet-Car. We report relative L2 error ( ↓ ) and Spearman’s rank correlation ( ↑ ). Top-2 results in each column are highlighted: best (orange) and second (blue).
Configuration
NS-Tracer-PwC
ShapeNet-Car
v↓
tracer ↓
Volume ↓
Surf ↓
CD↓
ρD↑
Advectra (Full)
0.1209
0.1607
0.0149
0.0716
0.0094
0.9941
Independent Warps
0.1548
0.2183
0.0174
0.0732
0.0101
0.9931
w/o Ordering
0.1446
0.1840
0.0169
0.0766
0.0113
0.9923
w/o Transport
0.1858
0.2435
0.0159
0.0742
0.0108
0.9928
w/o Both
0.2001
0.2525
0.0169
0.0797
0.0129
0.9918
Table 3 : Ablation study of Advectra. We evaluate Geometry-Aware Ordering, Latent Transport, and an independent-warp variant that uses separate source/target displacement maps instead of the shared transported update. Top-2 results per column are highlighted: best (orange) and second (blue).
Figure 2 : Visualizing the contrast between Eulerian and Lagrangian perspectives in Navier-Stokes flow. The Eulerian view treats the domain as a fixed grid, measuring fluid velocity at specific, stationary locations (cyan streamlines). In contrast, the Lagrangian view follows the individual trajectories of material parcels as they move through space. Here, a dense block of tracers (colored by initial position) acts as a set of “moving observers.” As t increases from 0 to 9 , these tracers are advected by the unsteady flow; their severe stretching and folding demonstrates how the Lagrangian frame of reference deforms over time, physically reconfiguring the spatial relationships between different regions of the fluid.
Model
AirCraft
Cp↓
ρ↓
U↓
V↓
W↓
Pressure ↓
GraphSAGE [ 24 ]
0.1442
0.0803
0.0148
0.2154
0.2395
0.0923
PointNet [ 25 ]
0.1472
0.0783
0.0147
0.2104
0.2528
0.0937
Graph U-Net [ 26 ]
0.1475
0.0784
0.0152
0.2225
0.2477
0.0947
MeshGraphNet [ 6 ]
0.1372
0.0682
0.0164
0.1835
0.2175
0.0893
GNO [ 8 ]
0.1556
0.0785
0.0167
0.2012
0.2318
0.1071
Table 4 : Quantitative results on the AirCraft benchmark (3D external aerodynamics). We compare against geometric baselines and neural operators. We report relative L2 error ( ↓ ). Top-2 results in each column are highlighted: best (orange) and second (blue).
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Model
COM ↓
Mass ↓
Raw MAE ↓
Ratio R↓
Advectra
0.0043
0.67%
0.0285
1.0099
Advectra (No Transport)
0.0071
0.93%
0.0419
1.0656
LaMO
0.0072
1.10%
0.0397
1.0973
Appendix
Table 5 : Transport diagnostics on NS-Tracer-PwC.
Figure 3 : Comparison of Raw Error versus Center-of-Mass (COM) Aligned Error across three randomly selected test trajectories. By translating the predicted tracer distribution by the displacement vector d=μgt−μpred before computing pointwise error, this diagnostic separates bulk COM misalignment from residual structural mismatch. Higher COM-aligned error indicates deformation, diffusion, or local misalignment that cannot be corrected by a global translation, while lower COM-aligned error indicates better preservation of tracer structure under this diagnostic.
Figure 4 : 1D horizontal cross-sections of the tracer concentration taken at the spatial peak for three randomly selected test trajectories. Advectra produces profiles that more closely match the ground truth in peak amplitude and spatial alignment. In contrast, the baseline models exhibit lower peak amplitudes, broader spatial spreads, and shifted centers, indicating stronger diffusion and advection lag.
Figure 5 : Qualitative comparison of predicted tracer transport across three randomly selected test trajectories. Each panel contrasts the Ground Truth state at tout against the predictions generated by the proposed Advectra architecture, the ablated Advectra (No Transport) variant, and the LaMO baseline. The visualizations illustrate the differing spatial distributions and structural deformations produced by each operator learning formulation.
Figure 6 : Distribution of consecutive centroid distances for randomly selected test samples. Geometry-aware ordering reduces both the average and median jump distance, indicating improved locality. At the same time, it preserves a small number of larger jumps, corresponding to structured long-range transitions necessary to capture global geometry.
Figure 7 : Latent sequence paths mapped to physical space for three randomly selected ShapeNet-Car test geometries. Without ordering, the sequence exhibits frequent long-range jumps between spatially discontinuous regions (e.g., front and rear of the geometry). With geometry-aware ordering, the paths become smoother and spatially coherent, traversing the domain through predominantly local transitions with only occasional long-range connections.
Figure 8 : Aggregation field visualization for three randomly selected ShapeNet-Car test geometries. Without ordering, the field exhibits long-range, crossing interactions spanning large portions of the domain, indicating spatially incoherent aggregation. With geometry-aware ordering, the field decomposes into localized, spatially coherent clusters with bounded extent, reflecting predominantly local interactions.
Sorting Interval
ShapeNet-Car
Volume ↓
Surf ↓
25
0.0158
0.0736
50
0.0150
0.0719
100
0.0149
0.0716
200
0.0153
0.0716
400
0.0156
0.0718
Appendix
Table 6 : Effect of ordering frequency on the ShapeNet-Car dataset. We report relative L2 errors on volume and surface fields. Geometry-aware ordering significantly improves performance compared to no ordering, while results remain stable across a wide range of sorting intervals (i.e., refresh intervals of the nearest-neighbor latent ordering).
Ordering Strategy
ShapeNet-Car
Volume ↓
Surf ↓
Greedy nearest-neighbor
0.0149
0.0716
Random-walk traversal
0.0159
0.0734
MST traversal
0.0166
0.0743
No sorting
0.0169
0.0766
Appendix
Table 7 : Effect of latent ordering strategy on the ShapeNet-Car dataset. We report relative L2 errors on volume and surface fields. Greedy nearest-neighbor ordering performs best among the tested traversal strategies.
Model
OOD Reynolds
OOD Angles
CL↓
ρL↑
CL↓
ρL↑
Simple MLP
0.6205
0.9578
0.4128
0.9572
GraphSAGE [ 24 ]
0.4333
0.9707
0.2538
0.9894
PointNet [ 25 ]
0.3836
0.9806
0.4425
0.9784
Graph U-Net [ 26 ]
0.4664
0.9645
0.3756
0.9816
MeshGraphNet [ 6 ]
1.7718
0.7631
0.6525
0.8927
Appendix
Table 8 : Generalization performance on Out-of-Distribution (OOD) tasks for the AirfRANS dataset. We evaluate extrapolation to unseen Reynolds numbers and angles of attack. ↓ indicates lower is better; ↑ indicates higher is better (Spearman’s rank correlation). Top-2 results in each column are highlighted: best (orange) and second (blue).
Model
Time (s/iter) ↓
GFLOPs / batch ↓
Transolver
0.1111
266.09
LaMO
0.1150
266.69
Advectra (Ours)
0.1209
258.06
Appendix
Table 9 : Runtime comparison on the Airfoil benchmark. We report average time per iteration and FLOPs per batch.
Figure 9 : Relative L2 error on the Airfoil benchmark as a function of token count. All methods improve with increasing token count and eventually saturate. Advectra reaches its saturated regime at a lower token count in this sweep, while the stronger 128 -slice setting benefits Transolver and LaMO.
Training Configuration
Model Configuration
Benchmark
Loss
Epochs
Init LR
Optimizer
Batch
Scheduler
L
Hr
D
G
Airfoil
Relative L2
500
10−3
AdamW
4
OneCycle
8
8
128
32
AirfRANS
Lv+Ls
400
10−3
Adam
1
OneCycle
8
8
128
32
ShapeNet-Car
Lv+0.5Ls
200
10−3
Adam
1
OneCycle
8
8
128
32
AirCraft
Lv+Ls
400
10−3
Adam
1
OneCycle
8
8
128
32
GCE-RT
Relative L2
500
10−3
AdamW
32
OneCycle
8
8
128
32
Appendix
Table 10 : Training and model configurations for Advectra. Baseline configurations follow official or best-reported settings when available, as described in Appendix I.4 . Relative L2 denotes field reconstruction error. Lv and Ls denote volume and surface losses. Here Hr denotes the number of parallel latent routing heads, i.e., independent anchor groups used for source/target assignments and feature projections. It is not a multi-head self-attention parameter; the latent sequence core is a bidirectional SSM.
OOD Reynolds
OOD Angles
Dataset
Range
Samples
Range
Samples
Training Set
[3×106,5×106]
500
[−2.5∘,12.5∘]
800
Test Set
[2×106,3×106]∪[5×106,6×106]
500
[−5∘,−2.5∘]∪[12.5∘,15∘]
200
Appendix
Table 11 : Settings for OOD generalization experiments on AirfRANS. The training and test sets contain disjoint ranges for Reynolds numbers and angles of attack.
GCE-RT
NS-Tracer-PwC
FNS-KF
Model
ρ
v
p
tracer
v
tracer
v
Best Baseline
0.0171±0.0008
0.0671±0.0012
0.0015±0.0001
0.0551±0.0014
0.1577±0.0021
0.2047±0.0068
0.1317±0.0010
(Transolver++)
(Transolver++)
(LinearNO)
(Transolver++)
(FNO)
(U-Net)
(FNO)
Advectra (Ours)
0.0149±0.0003
0.0595±0.0007
0.0014±0.0001
0.0488±0.0012
0.1209±0.0011
0.1607±0.0025
0.1202±0.0014
Appendix
Table 12 : Standard deviations for GCE-RT, NS-Tracer, and FNS-KF benchmarks. We report the mean ± standard deviation over three independent runs. All values are relative errors. The best baseline model for each task is indicated in parentheses.
Airfoil
ShapeNet-Car
AirfRANS
Model
(Rel. Error)
(Corr. %)
(Corr. %)
Best Baseline
0.0047±0.0001
99.28±0.02
99.92±0.01
(LaMO)
(SpiderSolver)
(LinearNO)
Advectra (Ours)
0.0041±0.0001
99.41±0.02
99.95±0.01
Appendix
Table 13 : Standard deviations for geometric and mesh-based benchmarks. We report the mean ± standard deviation over three independent runs. Spearman’s rank correlations for ShapeNet-Car and AirfRANS are presented as percentages (scaled by 100). Airfoil values are relative errors. The best baseline model for each task is indicated in parentheses.
Neural operators are fast, differentiable surrogates for physical simulation, but their accuracy often degrades when domain geometry, size, or operating conditions differ from training. Supervised adaptation can recover accuracy, but even a small target set requires costly high-fidelity simulations. We therefore ask how pretraining and transfer can be designed together to reduce this deployment cost. LatentDDM first pretrains a neural operator to predict fields on small subdomains. For a new setting, it freezes this operator and trains only a lightweight module that composes the local predictions. We evaluate our method on two complementary problems: steady Darcy flow, where long-range pressure coupling must extend across increasingly large porous domains, and unsteady incompressible flow around a pitching airfoil, where rollout errors compound as target pitching frequencies exceed the training range. Compared with the capacity-matched models that process the full domain at once, LatentDDM's error is 36-56% lower on larger Darcy domains after adaptation with 16 target simulations. It also improves 20-step field rollouts in fast-pitching airfoil flow, both zero-shot and after few-shot calibration. These results identify the co-designed local pretraining and composition-level transfer as a promising design principle for physical foundation models.
Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. However, we reveal that existing learnable projection mechanisms cannot ensure stable and balanced assignments from observation points to latent tokens, causing some latent tokens to be over-assigned while others remain underutilized. This limitation further restricts the design of hierarchical architectures, as assignment imbalance is continuously inherited and amplified across latent spaces, eventually causing severe token collapse in deeper spaces. To address these issues, we propose MoNo (Multiscale Optimal Transport Neural Operator), a progressive multiscale neural operator that efficiently solves PDEs on general geometries through stable latent-space construction. At its core is CoTAP (Cross-scale Optimal Transport Assignment and Projection), a novel latent-space construction method that formulates cross-space assignment between adjacent spaces as an entropy-regularized optimal transport problem, thereby constructing balanced bidirectional projections and stable latent spaces. CoTAP also ensures stable information transfer across multiple latent spaces, further enabling multiscale architectures on general geometries, which in turn support more efficient learning of long-range physical interactions. Extensive experiments demonstrate that MoNo outperforms existing state-of-the-art neural operators in both prediction performance and computational efficiency. Code is available at https://github.com/ZijiangY1116/MoNo.
Zijiang Yang, Xiaomeng Wu, Dongmei Fu
School of Automation and Electrical Engineering, University of Science and Technology Beijing · 2Beijing Engineering Research Center of Industrial Spectrum Imaging
We identify a mismatch between the physical role of transport fields in many PDEs and their usual role in neural operators: PDEs use them to select read coordinates, whereas neural operators typically treat them only as input values. We address this mismatch with the Green-Routed Neural Operator (GRNO), which uses the governing equation to determine where latent features are sampled. A parameter-free equation adapter evaluates the diagnostic relation and constructs a departure map whose values are the read coordinates. A multiscale encoder-decoder combines centered and routed reads of latent features to learn the complete finite-time update. Across five two- and three-dimensional PDE systems, GRNO achieves the lowest mean final relative L2 error on four under 40-step autoregressive evaluation and remains competitive on Keller-Segel. Fixed-weight route interventions reveal strong dependence on direction and spatial alignment in four systems, with weak dependence in Keller-Segel. In independently trained ablations, GRNO achieves lower mean errors than variants that supply the transport field only as an input feature, substitute a learned displacement for the equation-specified route, or apply the route with a spatial misalignment, across all five systems. It also substantially outperforms directly advecting the physical state and learning the remaining update, indicating that equation-specified read coordinates provide an effective structural prior for long-horizon PDE forecasting.
Chenhao Si, Ming Yan
School of Data Science The Chinese University of Hong Kong, Shenzhen Shenzhen, China