Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Authors: Alessio Borgi, Mario Severino, Fabrizio Silvestri, Pietro Liò
Organizations: Department of Computer Science and Technology, University of Cambridge · Department of Computer, Control and Management Engineering, Sapienza University of Rome · Department of Information Engineering, University of Padua
Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring vector features while preserving exact Euclidean equivariance. Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the edge transport itself. We characterize this transport theoretically, showing that when relative displacement is the only covariant geometric input, every linear O(n)-equivariant map decomposes into independent radial and tangential components, while learned covariant features enable richer feature-conditioned transformations. We also introduce controlled symmetry relaxation for systems with a preferred ambient direction, which may be prescribed or inferred from data while recovering full E(n)-equivariance when the directional pathway is inactive. Across particle dynamics, mesh-based simulation, point-cloud classification, and molecular property prediction, ESNN improves dynamics prediction, recovers the gravity axis when symmetry is broken, yields substantial gains on selected mesh tasks and long-horizon rollouts, and remains robust to unseen rotations. These results show that learning how geometric information is transported across edges offers a complementary route to expressive equivariant message passing without requiring higher-order representations.
Figures & tables
Figure 1: Edge-wise geometric transport in EGNN-style models, generic sheaf networks, and ESNNs. EGNN-style models couple invariant weights with relative displacements, while sheaf networks allow matrix-valued maps between local feature spaces. ESNN combines matrix-valued transport with exact equivariance through O(n) -covariant spatial actions and invariant channel mixing.
Figure 2: ESNN transport families. Identity leaves the vector representation unchanged; Diagonal applies an edge-dependent isotropic spatial scaling; Orthogonal learns a feature-conditioned spatial rotation; and Radial–Tangential acts independently on components parallel and orthogonal to the relative displacement.
Figure 3: Controlled symmetry relaxation through a preferred ambient direction. When all λg(ℓ)=0 , the directional pathway is inactive and ESNN retains full E(n) -equivariance. Activating the directional conditioning λg(ℓ)⟨rij,g⟩ reduces the guaranteed symmetry to Eg(n)=Og(n)⋉Rn , where Og(n)={Q∈O(n):Qg=g} . The layerwise relaxation coefficients are initialized at zero; g may be specified or learned, but is held fixed when the input geometry is transformed.
Table 1: N-body dynamics. Mean squared error (MSE) for charged and gravity-augmented prediction. Charged ESNN and gravity results are mean ± standard deviation over five runs; baseline uncertainties are reported as in the original sources. For gravity, maxℓ∣λg(ℓ)∣∥g∥2 measures directional-pathway activation and Ag alignment with the gravity axis; for Fixed , Ag=1 by construction. Lower MSE and higher Ag are better.
Method
z/z
z/SO(3)
SO(3)/SO(3)
ΔOOD↓
PointNet
85.9
19.6
74.7
66.3
RS-CNN
90.3
48.7
82.6
41.6
DGCNN
90.3
33.8
88.6
56.5
RI-Conv
86.5
86.4
86.4
0.1
GC-Conv
89.0
89.1
89.2
0.1
Luo et al. DGCNN
88.4
88.4
88.9
0.0
Table 2: ModelNet40 rotation generalization. Accuracy (%) under z/z , z/SO(3) , and SO(3)/SO(3) . The z/SO(3) result evaluates the same z -trained checkpoint under unseen SO(3) rotations. ΔOOD is the absolute z/z – z/SO(3) accuracy change in percentage points. Literature baselines follow Lippmann et al. (2025) ; EGNN and ESNN use the same evaluation protocol. Higher accuracy and lower ΔOOD are better.
CylinderFlow
DeformingPlate
Airfoil
Method
1-Step
50-Step
Full
1-Step
50-Step
Full
1-Step
50-Step
Full
MeshGraphNet
4.05±0.08
12.5±0.6
46.14±6.3
0.15±0.02
2.2±0.5
13.6±1.4
1141±68
1529±68
8033±662
EGNN
9.20±1.80
40.5±2.5
158.09±2.5
0.35±0.01
1.6±0.1
11.5±0.5
5001±134
18948±609
diverged
ESNN-Id
5.35±0.32
20.6±1.0
84.09±3.4
0.15±0.01
1.3±0.1
7.7±0.4
4744±154
7579±287
16893±740
ESNN-Diag
2.77±0.17
8.8±0.4
50.84±2.0
0.17±0.01
1.4±0.1
9.2±0.5
2413±49
4966±126
11014±321
ESNN-Ortho
2.56±0.15
8.5±0.4
55.42±2.2
0.13±0.01
1.1±0.1
8.1±0.4
1459±27
2807±66
8821±259
Table 3: Mesh-based physical dynamics. RMSE ( ×10−3 ) for one-step, 50-step, and full autoregressive rollouts on CylinderFlow , DeformingPlate , and Airfoil . Results are mean ± standard deviation over five matched runs. Best values are bold; Diverged denotes an unstable autoregressive rollout. Lower is better.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Spatial–channel factorization of ESNN transport. Each term acts on Vj∈Rn×cv along two independent axes: Sij(k) acts on the ambient spatial dimension by left multiplication, while Mij(k) acts on the vector channels by right multiplication. Summing the separable terms gives the canonical transport Ti←j . ESNN enforces equivariance by constraining the spatial operators to be O(n) -covariant and the channel maps to be invariant.
Figure 5: Charged-particle dynamics without and with a uniform gravitational field. Translation symmetry is preserved, while the ideal unclipped gravity–Coulomb dynamics retain only orthogonal transformations that preserve gtrue .
Figure 6: ModelNet40 graph construction and rotation protocols. Each centered point cloud is converted into a local k -nearest-neighbor graph and evaluated under the z/z , z/SO(3) , and SO(3)/SO(3) train/test regimes.
Figure 7: Mesh-based dynamics benchmarks: CylinderFlow , DeformingPlate , and Airfoil . Models are trained on one-step targets and evaluated autoregressively.
Method
α
Δϵ
ϵHOMO
ϵLUMO
μ
Cv
G
H
⟨R2⟩
U
U0
ZPVE
Units
ma03
meV
meV
meV
mD
mcalmol−1K−1
meV
meV
ma02
meV
meV
meV
Invariant models
Cormorant
85
61
34
38
38
26
20
21
961
21
22
2.03
NMP
92
69
43
38
30
40
19
17
180
20
20
1.50
DimeNet++ †
44
32.6
24.6
19.5
29.7
23
7.56
6.53
331
6.28
6.32
1.21
ComENet †
45
32.4
23.1
19.8
24.5
22
7.98
6.86
259
6.82
6.69
1.20
Appendix
Table 4: QM9 molecular-property prediction. Mean absolute error (MAE) on twelve invariant molecular properties. Baseline values and partition annotations follow Aykent and Xia (2025) ; † marks a different data partition. ESNN results are reported for the validation-selected model for each property, with the best ESNN result shown in bold . Units are given in the second header row. Lower is better.
Task
Model
Transport / training regime
C
L
V
Learning rate
Epochs
Charged N-body
ESNN
Identity
64
6
0
3.38×10−3
10000
ESNN
Diagonal
32
5
1
3.25×10−3
3000
ESNN
Orthogonal
32
7
1
1.92×10−3
3000
ESNN
Radial–Tangential
32
5
1
3.41×10−3
3000
Gravity N-body
ESNN
Diagonal
64
7
1
1.01×10−3
3000
ESNN
Orthogonal
32
7
1
1.24×10−3
3000
Appendix
Table 5: Selected N-body, gravity, and ModelNet40 configurations. C , L , and V denote scalar width, message-passing layers, and vector groups. For ModelNet40, z/z and z/SO(3) use the same z -trained checkpoint.