Cluster Attention Neural Operators for Solving Parametric Partial Differential Equations
Authors: Ming Zhong, Antonio Colanera, Gianluigi Rozza, Zhenya Yan
Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 100049, China. · State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing 100190, China. · International School for Advanced Studies (SISSA), Trieste 34136, Italy. · School of Mathematics and Information Science, Zhongyuan University of Technology, Zhengzhou 450007, China.
Traditional simulations of parametric partial differential equations (PDEs) rely on repetitive computations for each parameter, which makes high-fidelity design impractical. Neural operators address this issue by learning solution operators, accelerating parameter-space mapping by orders of magnitude. Recent Transformer-based neural operators attempt to capture global dependencies, but often at the cost of quadratic attention complexity. Transolver resolves this problem by projecting physical states into a reduced slice space for attention computation. Although fast, this projection sacrifices fine spatial information. Moreover, by operating in this reduced space with shared weights across attention heads, it may constrain the model's flexibility, thereby limiting its capacity to capture complex phenomena. To address these issues, we propose the Cluster Attention Neural Operator (CANO), which reformulates attention via a novel cross-attention mechanism that dynamically clusters queries while preserving full-resolution keys and values. This avoids slice compression loss and removes weight-sharing limits. At the same time, the model remains fast without losing global interactions. Empirically, CANO achieves state-of-the-art performance across canonical PDE benchmarks, covering fluid and solid dynamics (e.g., Navier-Stokes, Airfoil, Plasticity), irregular unstructured geometries (e.g., Pipe Turbulence, Composites), and long-term temporal rollouts. Across solid deformation and turbulent flow benchmarks, CANO achieves lower errors than baselines and exhibits strong geometric adaptability and temporal consistency.
Figures & tables
Figure 1: Comparison of the architecture of (a) Vanilla Self-Attention, (b) Transolver, and (c) CANO mechanisms.
Benchmark
Loss Function
Epochs
Batch
Cluster Number
L / H / d / rmlp
Standard benchmarks
Airfoil
Rel. L2
500
4
32
8 / 8 / 128 / 2
Pipe
Rel. L2
500
4
16
8 / 4 / 128 / 2
Plasticity
Rel. L2
500
8
64
8 / 8 / 128 / 2
Navier-Stokes
Rel. L2
500
2
32
8 / 8 / 256 / 1
Darcy
Rel. L2 + 0.1 L∇
500
4
128
8 / 8 / 128 / 2
Table 1: Unified training and model hyperparameters across benchmarks. All experiments utilize an initial learning rate of 10−3 . The model configuration is denoted as ( L / H / d / rmlp ), representing layer, number of attention head, embedding dimension and MLP ratio in the transformer blocks.
Test Case
Geometry
Dim.
Resolution
Dataset Split
Input Type
Train
Test
Airfoil [ 35 ]
Structured Mesh
2
221×51
1000
200
Boundary Shape
Pipe [ 35 ]
Structured Mesh
2
129×129
1000
200
Domain Shape
Plasticity [ 35 ]
Structured Mesh
2+1
101×31
900
80
External Force
Navier-Stokes [ 26 ]
Regular Grid
2+1
64×64
1000
200
Previous Vorticity
Darcy [ 26 ]
Regular Grid
2
85×85
1000
200
Coefficients
Table 2: Overview of standard PDE benchmarks. Details on geometry type, spatial dimension, discretization size, dataset splits, and input features.
Model
Structured Mesh
Regular Grid
Point Cloud
Airfoil
Pipe
Plasticity
Navier-Stokes
Darcy
Elasticity
(×10−2)
(×10−2)
(×10−2)
(×10−1)
(×10−2)
(×10−2)
Classic models
FNO [ 26 ]
/
/
/
1.56
1.08
/
DeepONet [ 27 ]
3.85
0.97
1.35
2.97
5.88
9.65
U-FNO [ 69 ]
2.69
0.56
0.39
2.23
1.83
2.39
Table 3: Performance comparison of neural operators on standard benchmarks (CANO vs. baselines). All values represent the relative L2 error. The best results are highlighted in bold , and the second-best are underlined . A slash (“/”) denotes benchmarks where the baseline is not applicable. Models marked with ∗ are reimplemented by us for fair comparison, while other results are taken from the original papers or the Transolver study.
Figure 2: Performance and distribution comparison: Relative L2 error of CANO and Transolver on six standard PDE datasets.
Figure 3: CANO predictions on representative PDE benchmarks. Steady-state physical fields and absolute errors for (a) Airfoil (structured mesh) and (b) Elasticity (point cloud). CANO accurately captures sharp gradients near complex boundaries. (c) Long-term temporal rollout predictions (from timestep +1 to +10) on the Navier-Stokes dataset. CANO effectively preserves intricate high-frequency eddies and structural integrity over time without severe error accumulation.
Figure 4: Visualization of the spatial assignments extracted from the final layer of the first head on the Airfoil dataset. (a) The corresponding ground truth, predictions and absolute errors. (b) The spatial activation maps of CANO’s cluster matrix C , showing allocations around discontinuities (e.g., shockwaves). (c) The slice allocation matrix of Transolver A , displaying relatively smooth regions.
Figure 5: Comparison of kinetic energy spectra ( 19 ) between CANO predictions and ground truth for the Navier-Stokes benchmark. Solid lines show the mean spectrum over 200 test samples; shaded regions indicate the mean ± 3 standard deviation.
Benchmark
Number of Parameters
Transolver
CANO (Ours)
Airfoil
3.07M
2.44M
Pipe
3.07M
2.49M
Plasticity
3.11M
2.51M
Navier-Stokes
12.28M
8.12M
Darcy
3.09M
2.42M
Table 4: Comparison of model parameters between Transolver and CANO across different standard cases.
Test Case
Geometry
Dim.
Nodes ( N )
Dataset Split
Input Type
Train
Test
Irregular Darcy [ 66 ]
Point Cloud
2
2290
1000
200
Coefficients
Pipe Turbulence [ 66 ]
Point Cloud
2
2673
300
100
Previous State
Composite [ 66 ]
Point Cloud
3
8232
400
100
Temperature
Table 5: Overview of irregular domain PDE benchmarks. Details on geometry type, spatial dimension, node discretization, dataset splits, and input features.
Model
Irregular Darcy
Pipe Turbulence
Composite
GraphSAGE [ 67 ]
6.73
23.60
20.90
DeepONet [ 27 ]
1.36
9.36
1.88
POD-DeepONet [ 31 ]
1.30
2.59
1.44
NORM [ 66 ]
1.05
1.01
1.00
HPM [ 72 ]
0.74
0.83
0.93
Transolver ∗ [ 58 ]
0.89
1.01
1.61
Table 6: Performance comparison of neural operators on Irregular Darcy, Pipe Turbulence, and Composite benchmarks (CANO vs. baselines). All values represent the relative L2 error (×10−4) . The best results are highlighted in bold , and the second-best are underlined .
Figure 6: Performance and qualitative distribution on irregular domains. (a) Raincloud plots comparing the relative L2 error distributions of CANO and Transolver. (b)-(d) Representative visual predictions of CANO for the (b) Irregular Darcy, (c) Pipe Turbulence, and (d) Composite benchmarks.
Test Case
Geometry
Nodes
Train Window
Test Window ( Tout )
Data Split
Input Type
Tin
Tout
Short
Long
Train
Test
Navier-Stokes [ 26 ]
Regular Grid
64×64
5
5
5
15
1000
200
ω
ICP Plasma [ 73 ]
Point Cloud
3424
5
5
5
50
499
150
ne,Te,v,T
Heat Transfer [ 73 ]
Point Cloud
4887 (Avg.)
5
5
5
50
740
132
u,v
Table 7: Overview of 2D long-time rollout benchmarks. Details on geometry type, node discretization, training/test window, dataset splits, and input features. Models are trained on short temporal windows ( Tout=5 ) but evaluated on significantly longer autoregressive rollouts (up to 15 or 50 steps) to test temporal stability.
Model
Navier-Stokes
ICP Plasma
Heat Transfer
+5
+15
+5
+50
+5
+50
Transolver
1.21
3.22
0.29
2.11
0.70
4.87
CANO
0.59
2.57
0.23
1.61
0.58
4.19
Table 8: Long-term predictive performance across different datasets. We evaluate the relative L2 error (×10−1) at different future time steps (e.g., +5, +15, +50). The best results are highlighted in bold .
Figure 7: Temporal error evolution and distribution analysis for long time rollouts. Panels (a) , (b) , and (c) correspond to the Navier-Stokes, ICP Plasma, and Heat Transfer datasets, respectively. Left: Per step relative L2 error accumulation curves, showing CANO’s slower error growth rate. Middle and Right: Raincloud plots comparing the error distributions at the short ( Tout=5 ) and long ( Tout=15 or 50 ) rollout stages.
Figure 8: Temporal error evolution on the ICP Plasma dataset. The visualization displays the absolute error of the predicted electron temperature Te for Transolver and CANO at timesteps +1, +25, and +50 compared to the ground truth.
Figure 9: Temporal error evolution on the Heat Transfer dataset. The visualization displays the absolute error of the predicted velocity magnitude ( u2+v2 ) for Transolver and CANO at timesteps +1, +25, and +50 compared to the ground truth.
Problem
Training Number
Transolver
CANO (ours)
Darcy Flow
200
2.17
1.80
400
0.98
0.72
600
0.71
0.54
800
0.67
0.46
1000
0.52
0.41
Navier-Stokes
200
3.76
3.09
Table 9: Relative L2 error comparison across different training sample sizes. All values for Darcy Flow are scaled by 10−2 , and Navier-Stokes by 10−1 . Best results are highlighted in bold .
Cluster Numeber ( Kc )
Darcy
Navier-Stokes
Pipe
Airfoil
Elasticity
Plasticity
8
5.38
9.05
3.83
4.96
5.08
0.88
16
4.83
7.67
3.32
4.77
4.78
0.78
32
4.73
6.58
3.55
4.29
4.72
0.81
64
4.36
6.36
3.48
4.88
4.81
0.71
128
4.08
5.60
3.82
4.38
4.62
0.77
Table 10: Relative L2 errors of CANO with different cluster numbers ( Kc ) across various PDE datasets. The best results for each dataset are highlighted in bold .
School of Automation and Electrical Engineering, University of Science and Technology Beijing · 2Beijing Engineering Research Center of Industrial Spectrum Imaging