Cross-Domain Pretraining for Steady-State Neural CFD Surrogates
Authors: Anthony Zhou, Amir Barati Farimani, Shirley Ho, Rudy Morel
Organizations: Carnegie Mellon University · New York University · Polymathic AI · Princeton University · Flatiron Institute, Center for Computational Astrophysics · Flatiron Institute, Center for Computational Mathematics · Flatiron Institute, Scientific Computing Core
Neural surrogates for computational fluid dynamics (CFD) have the potential to greatly enhance engineering innovation through accelerating simulation. However, the primary limitation for neural surrogates is the lack of generalization to geometries and applications beyond the training set, which is significant given the diversity of engineering scenarios. Currently, this is addressed by generating a new dataset for a specific application; however, this requires running costly numerical solvers. In this work, we take a step toward addressing this by studying neural surrogates trained across different geometries, boundary conditions, and fidelities. We find that cross-domain pretraining improves zero- and few-shot performance on held-out datasets relative to both training from scratch and transferring from domain-specific experts. In particular, finetuning a pretrained, cross-domain model can achieve 2-3x lower errors at the same sample size and use 8x fewer samples to achieve the same error, compared to training from scratch. This benefit is architecture agnostic and improves with model size and pretraining dataset diversity. Furthermore, we study how and why cross-domain pretraining works in CFD surrogates, and find that simply pooling steady-state datasets is both sufficient and effective. Given the high cost of generating CFD data, leveraging existing datasets through cross-domain pretraining will likely be a valuable strategy as future surrogates expand to tackle new problems and use cases.
Figures & tables
Figure 1: After pretraining a joint model across 6 CFD datasets, we evaluate its zero/few-shot performance on 4 held out datasets (see Table 1 ).
Name
Geometry
Closure
Solver
# Samples
# Cells
Re #
Inlet (m/s)
AoA ( ∘ )
DrivAerNet++
DrivAer Car
RANS κ-ω
OpenFOAM
8000
24M
8e6
30
-
DrivAerML
DrivAer Car
SA-DDES
OpenFOAM
500
160M
7.2e6
38.9
-
WindsorML
Windsor Body
WMLES
Volcano
355
280M
2.9e6
40
-
AhmedML ∗
Ahmed Body
SA-DDES
OpenFOAM
500
20M
7.7e5
1
-
SHIFT-Sub ∗
Submarine
RANS
Luminary
100
32M
3.2e7
5
-
Emmi Wing
Tapered Wings
RANS SA
OpenFOAM
30,000
3.3M
5-20e6
[150, 300]
[-10, 10]
Table 1: Datasets considered in this work, roughly split into automotive (top) and aerospace (bottom) applications, spanning a variety of geometries, fidelities and physical regimes. Datasets with an asterisk ∗ are held-out for downstream evaluation.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random init.
0.68
0.54
0.49
0.34
0.64
0.39
0.33
0.33
0.68
0.23
0.23
0.23
0.81
0.63
0.64
0.63
Domain Experts
DrivAerNet++
0.53
0.18
0.15
0.12
0.61
0.22
0.13
0.10
0.73
0.15
0.11
0.11
1.10
0.53
0.47
0.46
Emmi-Wing
0.73
0.21
0.21
0.19
0.66
0.24
0.19
0.16
0.61
0.14
0.12
0.12
1.10
0.52
0.48
0.49
Table 2: Validation error on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Figure 2: Zero-shot predictions ( Cp ) of pretrained models on a SHIFT-CCA geometry. Expert models can overfit to their pretraining sets, while the Joint model better predicts new geometries.
Figure 3: Validation loss vs. number of finetuning samples, plotted for the pretrained joint model and a model initialized from scratch. Two finetuning budgets are used (400/4000 steps). Note both axes are log scale.
Figure 4: Error of joint models on held-out datasets, after using S=0 or S=8 finetuning samples. Left: Joint models are pretrained with different numbers of pretraining datasets (2/4/6), with the model size held constant ( Np=32M ). Right: Joint models are pretrained at varying model sizes (9/32/121M), with the size of the pretraining dataset held constant ( Nd=6 ).
Model
DrivAerNet
Emmi-Wing
WindsorML
Double-Delta
DrivAerML
SuperWing
DrivAerNet
0.143
0.855
0.470
0.973
0.401
0.830
Emmi-Wing
0.809
0.030
0.617
1.239
0.830
0.742
WindsorML
0.678
0.747
0.067
0.956
0.701
0.736
Double-Delta
0.766
0.657
0.587
0.057
0.781
0.668
DrivAerML
0.421
0.786
0.458
0.966
0.048
0.803
SuperWing
0.973
0.648
0.705
1.255
1.006
0.047
Table 3: Validation error of expert and joint models across the pretraining datasets. Joint models are trained at different pretraining dataset sizes Nd=2,4,6 . Cells are colored by error magnitude.
Figure 8
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: The geometry (left), coefficient of pressure Cp (center), and x-component of the skin friction coefficient Cfx (right) visualized for a single sample in all datasets. From top to bottom: DrivAerNet, DrivAerML, WindsorML, AhmedML, SHIFT-Submarine, Emmi Wing, SuperWing, HiLiftAeroML, Double Delta, SHIFT-CCA.
Figure 7: The volumetric pressure (shown as the coefficient of pressure Cp ) and non-dimensional volumetric velocity magnitude ∣u∣/U∞ , plotted for a single sample in all datasets. From top to bottom: DrivAerNet, DrivAerML, WindsorML, AhmedML, SHIFT-Submarine, Emmi Wing, SuperWing, HiLiftAeroML, Double Delta, SHIFT-CCA.
SMART
SMART-IC
AB-UPT
Transolver++
Parameter
Value
Parameter
Value
Parameter
Value
Parameter
Value
Model size
32.0M
Model size
32.4M
Model size
36.1M
Model size
29.1M
Hidden dim
384
Hidden dim
384
Hidden dim
384
Hidden dim
384
Heads
8
Heads
8
Heads
8
Heads
8
Enc./dec. blocks
12
Enc./dec. blocks
8
Shared blocks
ppscscs
Layers
12
Num latents
4096
Num latents
4096
Supernodes
1024
Slices
32
Appendix
Table 5: Model hyperparameters. All models use a hidden width of 384 and are matched at roughly 30M parameters. The optimizer is also constant across all models (Adam ( β1=0.9 , β2=0.999 ), learning rate 10−4 , StepLR γ=0.99 every 1k steps)
Figure 8: Schematic of the SMART model variants. Cross-attention is routed so that the encoder and decoder branches can share weights (link symbols). Gray arrows indicate subsampling. Simulation parameters θ enter through Feature-wise Linear Modulation (FiLM) layers ( Perez et al., 2017 )
Figure 9: Schematics of AB-UPT and Transolver. Gray arrows indicate subsampling. For clarity, branches and weight sharing in AB-UPT, and slicing/unslicing in Transolver, are not shown. Simulation parameters θ enter through Adaptive LayerNorm (AdaLN) layers.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random init.
0.68
0.35
0.33
0.25
0.65
0.32
0.28
0.30
0.69
0.28
0.27
0.26
0.82
0.63
0.63
0.61
Domain Experts
DrivAerNet++
0.61
0.24
0.21
0.16
0.62
0.28
0.23
0.17
0.71
0.19
0.16
0.16
0.90
0.53
0.48
0.48
Emmi-Wing
0.74
0.24
0.22
0.19
0.68
0.29
0.24
0.22
0.70
0.20
0.18
0.17
1.02
0.60
0.54
0.53
Appendix
Table 6: Validation error for AB-UPT on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random Init.
0.68
0.36
0.34
0.30
0.64
0.31
0.29
0.27
0.68
0.22
0.21
0.20
0.81
0.61
0.59
0.58
Domain Experts
DrivAerNet++
0.59
0.27
0.27
0.19
0.57
0.28
0.24
0.22
0.79
0.21
0.19
0.18
0.97
0.56
0.46
0.51
Emmi-Wing
0.70
0.27
0.24
0.21
0.66
0.29
0.24
0.22
0.65
0.21
0.18
0.18
1.01
0.58
0.50
0.49
Appendix
Table 7: Validation error for Transolver++ on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Model
DrivAerNet
DrivAerML
Emmi-Wing
WindsorML
SuperWing
Double-Delta
Mean
Expert
SMART
0.142
0.047
0.030
0.065
0.047
0.055
0.064
AB-UPT
0.148
0.057
0.030
0.068
0.058
0.064
0.071
Transolver++
0.150
0.067
0.029
0.074
0.063
0.071
0.076
Joint
SMART
0.173
0.134
0.041
0.074
0.082
0.090
0.099
AB-UPT
0.177
0.148
0.041
0.080
0.089
0.097
0.105
Transolver++
0.184
0.166
0.044
0.097
0.107
0.124
0.120
Appendix
Table 8: SMART, AB-UPT, and Transolver++ models, evaluated on the pretraining datasets. Note that each cell in the Expert block is a different model (for example, the DrivAerNet SMART expert), while each row in the Joint block is a single model evaluated across all datasets.
Quantity
Dataset-specific
Global
Pressure coefficient
(Cp−μd(Cp))/σd(Cp)
Cp/fCp , fCp=0.33
Skin friction coefficient
(Cf−μd(Cf))/σd(Cf)
Cf/fCf , fCf=0.002
Velocity
(u∗−μd(u∗))/σd(u∗) , u∗=u/U∞
u′/fu′ , u′=(u−U∞)/U∞ , fu′=0.2
Simulation parameters
(θ−θmin)/(θmax−θmin)
(θ−θmin)/(θmax−θmin)
Geometry and queries
(G−cˉG)/sd , (x−cˉG)/sd
(G−cˉG)/sd , (x−cˉG)/sd
Estimated from simulation
μd,σd for every field
None
Appendix
Table 9: Comparison of dataset-specific and global normalization. The subscript d denotes a quantity computed separately for each dataset.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random Init.
0.82
0.42
0.40
0.33
0.82
0.38
0.31
0.30
0.85
0.22
0.21
0.21
0.90
0.65
0.65
0.64
Domain Experts
DrivAerNet++
0.53
0.15
0.13
0.12
0.89
0.19
0.12
0.09
0.89
0.13
0.11
0.10
0.94
0.55
0.47
0.46
Emmi-Wing
0.71
0.20
0.20
0.19
0.67
0.24
0.17
0.15
0.68
0.14
0.12
0.12
0.82
0.54
0.46
0.46
Appendix
Table 10: Validation error for SMART (with global normalization) on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Model
Normalization
DrivAerNet
DrivAerML
Emmi-Wing
WindsorML
SuperWing
Double-Delta
Mean
Expert
Dataset-specific
0.142
0.047
0.030
0.067
0.048
0.057
0.065
Global
0.142
0.046
0.030
0.067
0.047
0.058
0.065
Joint
Dataset-specific
0.173
0.134
0.043
0.074
0.082
0.090
0.099
Global
0.172
0.132
0.041
0.074
0.079
0.091
0.098
Appendix
Table 11: Comparison of dataset-specific and global normalization during pretraining. Losses are computed in the de-normalized space (raw Cp,Cf,u∗ ). Note that each cell in the Expert block is a different model, while each row in the Joint block is a single model evaluated across all datasets.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random Init.
0.77
0.57
0.54
0.40
0.83
0.53
0.49
0.47
0.87
0.53
0.45
0.46
0.87
0.76
0.75
0.76
Domain Experts
DrivAerNet++
0.57
0.17
0.14
0.12
0.64
0.31
0.23
0.19
0.88
0.34
0.26
0.24
0.97
0.65
0.57
0.49
Emmi-Wing
0.81
0.22
0.21
0.19
0.77
0.38
0.31
0.27
0.77
0.28
0.26
0.26
0.98
0.62
0.54
0.57
Appendix
Table 12: Validation error for SMART (evaluated on simulation mesh) on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Model
DrivAerNet
Emmi-Wing
WindsorML
Double-Delta
DrivAerML
SuperWing
Joint (Base)
0.174
0.043
0.075
0.091
0.134
0.082
Joint (121M)
0.150
0.039
0.072
0.070
0.074
0.074
Joint (6xSteps)
0.150
0.029
0.069
0.063
0.073
0.041
Appendix
Table 13: Validation error of the joint model across the pretraining datasets, either with an increased model size (121M) or increased gradient budget (6xSteps).
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
Joint (Base)
0.31
0.12
0.10
0.09
0.52
0.14
0.10
0.09
0.52
0.11
0.10
0.10
0.90
0.46
0.34
0.33
Joint (121M)
0.31
0.12
0.09
0.08
0.53
0.16
0.11
0.08
0.55
0.09
0.08
0.08
0.92
0.45
0.32
0.28
Joint (6xSteps)
0.32
0.12
0.10
0.08
0.61
0.16
0.10
0.09
0.55
0.10
0.09
0.09
1.11
0.44
0.32
0.30
Appendix
Table 14: Validation error on four held-out target datasets. Each row is a different joint model, either with increased model size (121M) or increased gradient budget (6xSteps). Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation.
Model
DrivAerNet
Emmi-Wing
WindsorML
Double-Delta
DrivAerML
SuperWing
Joint (Base)
0.174
0.043
0.075
0.091
0.134
0.082
Joint (IC)
0.176
0.075
0.076
0.098
0.144
0.098
Appendix
Table 15: Pretraining validation error of the joint base model (SMART) and the in-context model (SMART-IC) on each dataset. SMART-IC additionally receives a solved context simulation from the same dataset as the query. Providing context does not reduce the error on any dataset.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random init.
0.82
0.74
0.65
0.25
0.88
0.53
0.39
0.37
0.97
0.18
0.16
0.16
0.96
0.68
0.69
0.68
Domain Experts
DrivAerNet++
0.65
0.21
0.17
0.14
0.87
0.36
0.20
0.13
0.96
0.13
0.08
0.07
1.08
0.59
0.53
0.52
Emmi-Wing
0.85
0.24
0.22
0.20
0.90
0.40
0.30
0.22
0.89
0.12
0.10
0.10
1.12
0.58
0.51
0.52
Appendix
Table 16: Validation error in surface pressure coefficient Cp on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random init.
0.65
0.61
0.60
0.49
0.41
0.35
0.34
0.32
0.41
0.30
0.31
0.31
0.77
0.73
0.74
0.73
Domain Experts
DrivAerNet++
0.64
0.17
0.14
0.12
0.45
0.19
0.13
0.12
0.44
0.20
0.18
0.18
0.86
0.62
0.56
0.55
Emmi-Wing
0.75
0.19
0.19
0.18
0.49
0.19
0.16
0.15
0.39
0.19
0.18
0.18
1.39
0.63
0.59
0.61
Appendix
Table 17: Validation error in surface skin-friction coefficient Cf on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random init.
0.98
0.57
0.47
0.36
1.00
0.42
0.33
0.36
0.99
0.21
0.21
0.21
0.98
0.62
0.63
0.63
Domain Experts
DrivAerNet++
0.63
0.24
0.21
0.18
0.85
0.24
0.13
0.09
1.14
0.14
0.09
0.08
1.48
0.53
0.46
0.46
Emmi-Wing
1.01
0.28
0.30
0.27
1.01
0.29
0.21
0.18
0.91
0.13
0.10
0.10
1.13
0.51
0.46
0.47
Appendix
Table 18: Validation error in volume pressure coefficient Cp on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Held-out Datasets:
AhmedML
SHIFT-Sub
SHIFT-CCA
HiLiftAeroML
# FT Samples:
0
2
4
8
0
2
4
8
0
2
4
8
0
2
4
8
No Pretraining
Random init.
0.26
0.26
0.25
0.25
0.26
0.25
0.25
0.25
0.35
0.24
0.25
0.25
0.53
0.49
0.50
0.49
Domain Experts
DrivAerNet++
0.19
0.09
0.07
0.06
0.27
0.08
0.07
0.06
0.39
0.12
0.11
0.11
0.97
0.38
0.33
0.33
Emmi-Wing
0.29
0.12
0.13
0.12
0.24
0.09
0.09
0.08
0.25
0.12
0.11
0.12
0.77
0.37
0.35
0.36
Appendix
Table 19: Validation error in volume velocity u∗ on four held-out target datasets. Rows indicate the model initialization; single-domain experts are identified by their pretraining dataset. Columns give the number of fine-tuning samples, with 0 denoting zero-shot evaluation. Green and red indicate the best and worst single-domain experts, respectively, and bold indicates the lowest overall error.
Figure 10: Model predictions of surface Cp on AhmedML.
Figure 11: Model predictions of surface Cfx on AhmedML.
Figure 12: Model predictions of volume Cp on AhmedML.
Figure 13: Model predictions of volume ∣u∗∣ on AhmedML.
Figure 14: Model predictions of surface Cp on SHIFT-Submarine.
Figure 15: Model predictions of surface Cfx on SHIFT-Submarine.
Figure 16: Model predictions of volume Cp on SHIFT-Submarine.
Figure 17: Model predictions of volume ∣u∗∣ on SHIFT-Submarine.
Figure 18: Model predictions of surface Cp on SHIFT-CCA.
Figure 19: Model predictions of surface Cfx on SHIFT-CCA.
Figure 20: Model predictions of volume Cp on SHIFT-CCA.
Figure 21: Model predictions of volume ∣u∗∣ on SHIFT-CCA
Figure 22: Model predictions of surface Cp on HiLiftAeroML.
Figure 23: Model predictions of surface Cfx on HiLiftAeroML.
Figure 24: Model predictions of volume Cp on HiLiftAeroML.
Figure 25: Model predictions of volume ∣u∗∣ on HiLiftAeroML
Pretraining a neural PDE surrogate can reduce the amount of new CFD data needed when geometry or modeled physics changes. However, it remains unclear how different components of distribution shift affect this benefit. We pretrain a surrogate on 254,909 RANS solutions from one airfoil family and fine-tune it on a new family under two target settings with matched freestream ranges: the same Spalart-Allmaras (SA) modeling and SA with added eN transition modeling. At N=1000, the pretrained model matches the accuracy of a model trained from scratch on 3.25× as many samples for the same-SA target, but 2.58× as many for the transition-modeled target. By N=5000, this ordering reverses (1.56× versus 1.86×). At N=1000, sampling more distinct airfoils lowers error on both targets, but only for the same-SA target is the gain increase larger than the observed draw-to-draw variation (3.3× to 4.0×). These results show that pretraining value depends jointly on target-data budget, target-data coverage, and whether source and target differ in modeled physics.
Pochinapeddi Sai Bhargav, Nithin Somasekharan, Rohit Sunil Kanchi +2
Rensselaer Polytechnic Institute Troy, NY 12180 · University of Tennessee Knoxville, TN 37996
Neural surrogate models for computational fluid dynamics (CFD) are typically trained as forward operators that map explicit problem specifications, such as geometry and boundary conditions, to solution fields. This ties the model to the conditioning variables seen during training and limits reuse under boundary-condition shifts or local geometry changes. We propose to reformulate steady CFD inference as an inpainting problem: instead of training on explicit boundary conditions, we learn a self-supervised prior over velocity fields and impose boundary constraints only during inference by fixing known regions such as inlet, outlet or unchanged regions from previous simulations. To scale this idea to large 3D meshes, we introduce a local neighbourhood tokeniser that represents high-resolution velocity fields as compact spatial latent tokens and train latent flow-matching and masked-autoencoder models on these tokens. On intracranial aneurysm hemodynamics, our method reconstructs full velocity fields from sparse boundary context, outperforms supervised neural surrogates under boundary-condition and dataset shift and enables local geometry editing by reusing unchanged simulation context. These results suggest that viewing CFD inference as context-conditioned inpainting can turn neural surrogates from task-specific predictors into reusable flow priors.
Jonas Weidner, Yeray Martin-Ruisanchez, Daniel Rueckert +2
AI for Image-Guided Diagnosis and Therapy, Technical University of Munich · Munich Center for Machine Learning (MCML) · AI in Healthcare and Medicine, Technical University of Munich +1
Neural surrogates can accelerate computational fluid dynamics (CFD) simulations by orders of magnitude, but practical deployment in engineering and healthcare applications requires architectures that scale to high-resolution meshes and learn effectively from limited data. Explicit equivariance offers a principled inductive bias, yet its accuracy benefits may depend on the prediction task and the distribution of anatomical orientations. We investigate this dependence across three hemodynamic benchmarks with different degrees of natural canonical alignment. To support this study, we introduce the Anchored-Branched Geometric Algebra Transformer (AB-GATr), an E(3)-equivariant surrogate that efficiently predicts coupled surface and volume quantities. Across these benchmarks, AB-GATr consistently outperforms the evaluated non-equivariant models, including variants trained with rotational augmentation, while achieving accuracy competitive with E(3)-equivariant LaB-GATr at substantially lower training cost. In comparison, rotational augmentation provides inconsistent benefits across architectures and can reduce accuracy. A controlled experiment on ShapeNet-Car shows that strong canonical alignment can favor non-equivariant models, but their accuracy generally deteriorates as training orientations broaden and can decline sharply under broader test rotations. We further investigate these patterns using extended symmetry-breaking diagnostics and probes of the predictive information associated with canonical alignment across all benchmarks. Together, these results support explicit equivariance for the evaluated hemodynamic tasks with natural orientation variation, while showing that its accuracy benefits depend on the task and orientation distribution.
Patryk Rygiel, Julian Suk, Kak Khee Yeung +2
Department of Applied Mathematics Technical Medical Centre Cardiovascular Health Technology Centre University of Twente · Department of Computer Science Munich Center for Machine Learning Technical University of Munich · Department of Surgery Amsterdam UMC, Location University of Amsterdam Atherosclerosis & Aortic diseases Amsterdam Cardiovascular Sciences