Cluster Attention Neural Operators for Solving Parametric Partial Differential Equations
Authors: Ming Zhong, Antonio Colanera, Gianluigi Rozza, Zhenya Yan
Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 100049, China. · State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing 100190, China. · International School for Advanced Studies (SISSA), Trieste 34136, Italy. · School of Mathematics and Information Science, Zhongyuan University of Technology, Zhengzhou 450007, China.
Traditional simulations of parametric partial differential equations (PDEs) rely on repetitive computations for each parameter, which makes high-fidelity design impractical. Neural operators address this issue by learning solution operators, accelerating parameter-space mapping by orders of magnitude. Recent Transformer-based neural operators attempt to capture global dependencies, but often at the cost of quadratic attention complexity. Transolver resolves this problem by projecting physical states into a reduced slice space for attention computation. Although fast, this projection sacrifices fine spatial information. Moreover, by operating in this reduced space with shared weights across attention heads, it may constrain the model's flexibility, thereby limiting its capacity to capture complex phenomena. To address these issues, we propose the Cluster Attention Neural Operator (CANO), which reformulates attention via a novel cross-attention mechanism that dynamically clusters queries while preserving full-resolution keys and values. This avoids slice compression loss and removes weight-sharing limits. At the same time, the model remains fast without losing global interactions. Empirically, CANO achieves state-of-the-art performance across canonical PDE benchmarks, covering fluid and solid dynamics (e.g., Navier-Stokes, Airfoil, Plasticity), irregular unstructured geometries (e.g., Pipe Turbulence, Composites), and long-term temporal rollouts. Across solid deformation and turbulent flow benchmarks, CANO achieves lower errors than baselines and exhibits strong geometric adaptability and temporal consistency.
Figures & tables
Figure 1: Comparison of the architecture of (a) Vanilla Self-Attention, (b) Transolver, and (c) CANO mechanisms.
Benchmark
Loss Function
Epochs
Batch
Cluster Number
L / H / d / rmlp
Standard benchmarks
Airfoil
Rel. L2
500
4
32
8 / 8 / 128 / 2
Pipe
Rel. L2
500
4
16
8 / 4 / 128 / 2
Plasticity
Rel. L2
500
8
64
8 / 8 / 128 / 2
Navier-Stokes
Rel. L2
500
2
32
8 / 8 / 256 / 1
Darcy
Rel. L2 + 0.1 L∇
500
4
128
8 / 8 / 128 / 2
Table 1: Unified training and model hyperparameters across benchmarks. All experiments utilize an initial learning rate of 10−3 . The model configuration is denoted as ( L / H / d / rmlp ), representing layer, number of attention head, embedding dimension and MLP ratio in the transformer blocks.
Test Case
Geometry
Dim.
Resolution
Dataset Split
Input Type
Train
Test
Airfoil [ 35 ]
Structured Mesh
2
221×51
1000
200
Boundary Shape
Pipe [ 35 ]
Structured Mesh
2
129×129
1000
200
Domain Shape
Plasticity [ 35 ]
Structured Mesh
2+1
101×31
900
80
External Force
Navier-Stokes [ 26 ]
Regular Grid
2+1
64×64
1000
200
Previous Vorticity
Darcy [ 26 ]
Regular Grid
2
85×85
1000
200
Coefficients
Table 2: Overview of standard PDE benchmarks. Details on geometry type, spatial dimension, discretization size, dataset splits, and input features.
Model
Structured Mesh
Regular Grid
Point Cloud
Airfoil
Pipe
Plasticity
Navier-Stokes
Darcy
Elasticity
(×10−2)
(×10−2)
(×10−2)
(×10−1)
(×10−2)
(×10−2)
Classic models
FNO [ 26 ]
/
/
/
1.56
1.08
/
DeepONet [ 27 ]
3.85
0.97
1.35
2.97
5.88
9.65
U-FNO [ 69 ]
2.69
0.56
0.39
2.23
1.83
2.39
Table 3: Performance comparison of neural operators on standard benchmarks (CANO vs. baselines). All values represent the relative L2 error. The best results are highlighted in bold , and the second-best are underlined . A slash (“/”) denotes benchmarks where the baseline is not applicable. Models marked with ∗ are reimplemented by us for fair comparison, while other results are taken from the original papers or the Transolver study.
Figure 2: Performance and distribution comparison: Relative L2 error of CANO and Transolver on six standard PDE datasets.
Figure 3: CANO predictions on representative PDE benchmarks. Steady-state physical fields and absolute errors for (a) Airfoil (structured mesh) and (b) Elasticity (point cloud). CANO accurately captures sharp gradients near complex boundaries. (c) Long-term temporal rollout predictions (from timestep +1 to +10) on the Navier-Stokes dataset. CANO effectively preserves intricate high-frequency eddies and structural integrity over time without severe error accumulation.
Figure 4: Visualization of the spatial assignments extracted from the final layer of the first head on the Airfoil dataset. (a) The corresponding ground truth, predictions and absolute errors. (b) The spatial activation maps of CANO’s cluster matrix C , showing allocations around discontinuities (e.g., shockwaves). (c) The slice allocation matrix of Transolver A , displaying relatively smooth regions.
Figure 5: Comparison of kinetic energy spectra ( 19 ) between CANO predictions and ground truth for the Navier-Stokes benchmark. Solid lines show the mean spectrum over 200 test samples; shaded regions indicate the mean ± 3 standard deviation.
Benchmark
Number of Parameters
Transolver
CANO (Ours)
Airfoil
3.07M
2.44M
Pipe
3.07M
2.49M
Plasticity
3.11M
2.51M
Navier-Stokes
12.28M
8.12M
Darcy
3.09M
2.42M
Table 4: Comparison of model parameters between Transolver and CANO across different standard cases.
Test Case
Geometry
Dim.
Nodes ( N )
Dataset Split
Input Type
Train
Test
Irregular Darcy [ 66 ]
Point Cloud
2
2290
1000
200
Coefficients
Pipe Turbulence [ 66 ]
Point Cloud
2
2673
300
100
Previous State
Composite [ 66 ]
Point Cloud
3
8232
400
100
Temperature
Table 5: Overview of irregular domain PDE benchmarks. Details on geometry type, spatial dimension, node discretization, dataset splits, and input features.
Model
Irregular Darcy
Pipe Turbulence
Composite
GraphSAGE [ 67 ]
6.73
23.60
20.90
DeepONet [ 27 ]
1.36
9.36
1.88
POD-DeepONet [ 31 ]
1.30
2.59
1.44
NORM [ 66 ]
1.05
1.01
1.00
HPM [ 72 ]
0.74
0.83
0.93
Transolver ∗ [ 58 ]
0.89
1.01
1.61
Table 6: Performance comparison of neural operators on Irregular Darcy, Pipe Turbulence, and Composite benchmarks (CANO vs. baselines). All values represent the relative L2 error (×10−4) . The best results are highlighted in bold , and the second-best are underlined .
Figure 6: Performance and qualitative distribution on irregular domains. (a) Raincloud plots comparing the relative L2 error distributions of CANO and Transolver. (b)-(d) Representative visual predictions of CANO for the (b) Irregular Darcy, (c) Pipe Turbulence, and (d) Composite benchmarks.
Test Case
Geometry
Nodes
Train Window
Test Window ( Tout )
Data Split
Input Type
Tin
Tout
Short
Long
Train
Test
Navier-Stokes [ 26 ]
Regular Grid
64×64
5
5
5
15
1000
200
ω
ICP Plasma [ 73 ]
Point Cloud
3424
5
5
5
50
499
150
ne,Te,v,T
Heat Transfer [ 73 ]
Point Cloud
4887 (Avg.)
5
5
5
50
740
132
u,v
Table 7: Overview of 2D long-time rollout benchmarks. Details on geometry type, node discretization, training/test window, dataset splits, and input features. Models are trained on short temporal windows ( Tout=5 ) but evaluated on significantly longer autoregressive rollouts (up to 15 or 50 steps) to test temporal stability.
Model
Navier-Stokes
ICP Plasma
Heat Transfer
+5
+15
+5
+50
+5
+50
Transolver
1.21
3.22
0.29
2.11
0.70
4.87
CANO
0.59
2.57
0.23
1.61
0.58
4.19
Table 8: Long-term predictive performance across different datasets. We evaluate the relative L2 error (×10−1) at different future time steps (e.g., +5, +15, +50). The best results are highlighted in bold .
Figure 7: Temporal error evolution and distribution analysis for long time rollouts. Panels (a) , (b) , and (c) correspond to the Navier-Stokes, ICP Plasma, and Heat Transfer datasets, respectively. Left: Per step relative L2 error accumulation curves, showing CANO’s slower error growth rate. Middle and Right: Raincloud plots comparing the error distributions at the short ( Tout=5 ) and long ( Tout=15 or 50 ) rollout stages.
Figure 8: Temporal error evolution on the ICP Plasma dataset. The visualization displays the absolute error of the predicted electron temperature Te for Transolver and CANO at timesteps +1, +25, and +50 compared to the ground truth.
Figure 9: Temporal error evolution on the Heat Transfer dataset. The visualization displays the absolute error of the predicted velocity magnitude ( u2+v2 ) for Transolver and CANO at timesteps +1, +25, and +50 compared to the ground truth.
Problem
Training Number
Transolver
CANO (ours)
Darcy Flow
200
2.17
1.80
400
0.98
0.72
600
0.71
0.54
800
0.67
0.46
1000
0.52
0.41
Navier-Stokes
200
3.76
3.09
Table 9: Relative L2 error comparison across different training sample sizes. All values for Darcy Flow are scaled by 10−2 , and Navier-Stokes by 10−1 . Best results are highlighted in bold .
Cluster Numeber ( Kc )
Darcy
Navier-Stokes
Pipe
Airfoil
Elasticity
Plasticity
8
5.38
9.05
3.83
4.96
5.08
0.88
16
4.83
7.67
3.32
4.77
4.78
0.78
32
4.73
6.58
3.55
4.29
4.72
0.81
64
4.36
6.36
3.48
4.88
4.81
0.71
128
4.08
5.60
3.82
4.38
4.62
0.77
Table 10: Relative L2 errors of CANO with different cluster numbers ( Kc ) across various PDE datasets. The best results for each dataset are highlighted in bold .
Neural operators have emerged as powerful data-driven solvers for PDEs, offering substantial acceleration over classical numerical methods. However, existing transformer-based operators still face critical challenges when modeling PDEs on complex geometries: directly processing over massive mesh points is computationally expensive, while operating in raw discretization coordinates may obscure the intrinsic geometry where physical interactions are more naturally expressed. To address these limitations, we introduce the Charted Axial Transformer Operator (CATO), a geometry-adaptive and derivative-aware neural operator for PDEs on general geometries. Instead of applying attention directly in the physical coordinate system, CATO learns a continuous latent chart that maps mesh coordinates into a learned chart space, where chart-conditioned axial attention efficiently captures long-range dependencies with reduced computational cost. In addition, CATO introduces a derivative-aware physics loss for steady-state PDEs that jointly supervises solution values, mesh-consistent gradients, and an auxiliary flux-like field, improving physical fidelity and reducing oversmoothing. We further provide a theoretical approximation result showing that, under a favorable chart, charted axial attention can represent low-rank axial solution operators with controlled error, and that small chart perturbations induce bounded approximation degradation. CATO achieves the best performance across all evaluated datasets, yielding an average improvement of approximately 26.76% over the strongest competing baselines while reducing the number of parameters by 81.98%. These results highlight the effectiveness of learning geometry-adaptive charts and derivative-aware physical supervision for accurate and efficient PDE operator learning.
Neural operators have become a common approach for learning PDE solution maps and accelerating numerical simulations. Transformer-based neural operators are of particular interest, since attention can learn long-range dependencies in the computational domain. However, standard attention has two major limitations when applied to PDEs: it scales quadratically with the number of computational nodes, and it lacks an explicit bias toward local interactions. To address these issues, we introduce Local Linear Transformer (LLT) for PDE operator learning. The architecture combines linear global attention with local spatial mixing, and incorporates coordinate and geometry information. We evaluate LLT on several PDE problems, including elasticity, plasticity, airfoil flow, pipe flow, and Darcy flow. The reference data for these problems span finite-element, finite-volume, and finite-difference discretizations on structured and unstructured meshes. Compared with other neural-operator and transformer baselines from prior studies, LLT achieves competitive or lower relative L2 error across these problems. On matched structured discretizations, wall-clock time per training iteration is reduced by factors of 1.8 to 2.5 relative to Transolver. We also scale the approach and apply it to a three-dimensional car aerodynamics dataset with 32,186 unstructured mesh points per sample. Together, these results indicate that LLT provides an accurate and computationally efficient operator for PDE problems across discretizations, mesh types, and problem settings.
Oded Ovadia, Eli Turkel
School of Mathematical Sciences, Tel Aviv University, Tel Aviv, Israel
Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. However, we reveal that existing learnable projection mechanisms cannot ensure stable and balanced assignments from observation points to latent tokens, causing some latent tokens to be over-assigned while others remain underutilized. This limitation further restricts the design of hierarchical architectures, as assignment imbalance is continuously inherited and amplified across latent spaces, eventually causing severe token collapse in deeper spaces. To address these issues, we propose MoNo (Multiscale Optimal Transport Neural Operator), a progressive multiscale neural operator that efficiently solves PDEs on general geometries through stable latent-space construction. At its core is CoTAP (Cross-scale Optimal Transport Assignment and Projection), a novel latent-space construction method that formulates cross-space assignment between adjacent spaces as an entropy-regularized optimal transport problem, thereby constructing balanced bidirectional projections and stable latent spaces. CoTAP also ensures stable information transfer across multiple latent spaces, further enabling multiscale architectures on general geometries, which in turn support more efficient learning of long-range physical interactions. Extensive experiments demonstrate that MoNo outperforms existing state-of-the-art neural operators in both prediction performance and computational efficiency. Code is available at https://github.com/ZijiangY1116/MoNo.
Zijiang Yang, Xiaomeng Wu, Dongmei Fu
School of Automation and Electrical Engineering, University of Science and Technology Beijing · 2Beijing Engineering Research Center of Industrial Spectrum Imaging