Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Authors: Alessio Borgi, Mario Severino, Fabrizio Silvestri, Pietro Liò
Organizations: Department of Computer Science and Technology, University of Cambridge · Department of Computer, Control and Management Engineering, Sapienza University of Rome · Department of Information Engineering, University of Padua
Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring vector features while preserving exact Euclidean equivariance. Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the edge transport itself. We characterize this transport theoretically, showing that when relative displacement is the only covariant geometric input, every linear O(n)-equivariant map decomposes into independent radial and tangential components, while learned covariant features enable richer feature-conditioned transformations. We also introduce controlled symmetry relaxation for systems with a preferred ambient direction, which may be prescribed or inferred from data while recovering full E(n)-equivariance when the directional pathway is inactive. Across particle dynamics, mesh-based simulation, point-cloud classification, and molecular property prediction, ESNN improves dynamics prediction, recovers the gravity axis when symmetry is broken, yields substantial gains on selected mesh tasks and long-horizon rollouts, and remains robust to unseen rotations. These results show that learning how geometric information is transported across edges offers a complementary route to expressive equivariant message passing without requiring higher-order representations.
Figures & tables
Figure 1: Edge-wise geometric transport in EGNN-style models, generic sheaf networks, and ESNNs. EGNN-style models couple invariant weights with relative displacements, while sheaf networks allow matrix-valued maps between local feature spaces. ESNN combines matrix-valued transport with exact equivariance through O(n) -covariant spatial actions and invariant channel mixing.
Figure 2: ESNN transport families. Identity leaves the vector representation unchanged; Diagonal applies an edge-dependent isotropic spatial scaling; Orthogonal learns a feature-conditioned spatial rotation; and Radial–Tangential acts independently on components parallel and orthogonal to the relative displacement.
Figure 3: Controlled symmetry relaxation through a preferred ambient direction. When all λg(ℓ)=0 , the directional pathway is inactive and ESNN retains full E(n) -equivariance. Activating the directional conditioning λg(ℓ)⟨rij,g⟩ reduces the guaranteed symmetry to Eg(n)=Og(n)⋉Rn , where Og(n)={Q∈O(n):Qg=g} . The layerwise relaxation coefficients are initialized at zero; g may be specified or learned, but is held fixed when the input geometry is transformed.
Table 1: N-body dynamics. Mean squared error (MSE) for charged and gravity-augmented prediction. Charged ESNN and gravity results are mean ± standard deviation over five runs; baseline uncertainties are reported as in the original sources. For gravity, maxℓ∣λg(ℓ)∣∥g∥2 measures directional-pathway activation and Ag alignment with the gravity axis; for Fixed , Ag=1 by construction. Lower MSE and higher Ag are better.
Method
z/z
z/SO(3)
SO(3)/SO(3)
ΔOOD↓
PointNet
85.9
19.6
74.7
66.3
RS-CNN
90.3
48.7
82.6
41.6
DGCNN
90.3
33.8
88.6
56.5
RI-Conv
86.5
86.4
86.4
0.1
GC-Conv
89.0
89.1
89.2
0.1
Luo et al. DGCNN
88.4
88.4
88.9
0.0
Table 2: ModelNet40 rotation generalization. Accuracy (%) under z/z , z/SO(3) , and SO(3)/SO(3) . The z/SO(3) result evaluates the same z -trained checkpoint under unseen SO(3) rotations. ΔOOD is the absolute z/z – z/SO(3) accuracy change in percentage points. Literature baselines follow Lippmann et al. (2025) ; EGNN and ESNN use the same evaluation protocol. Higher accuracy and lower ΔOOD are better.
CylinderFlow
DeformingPlate
Airfoil
Method
1-Step
50-Step
Full
1-Step
50-Step
Full
1-Step
50-Step
Full
MeshGraphNet
4.05±0.08
12.5±0.6
46.14±6.3
0.15±0.02
2.2±0.5
13.6±1.4
1141±68
1529±68
8033±662
EGNN
9.20±1.80
40.5±2.5
158.09±2.5
0.35±0.01
1.6±0.1
11.5±0.5
5001±134
18948±609
diverged
ESNN-Id
5.35±0.32
20.6±1.0
84.09±3.4
0.15±0.01
1.3±0.1
7.7±0.4
4744±154
7579±287
16893±740
ESNN-Diag
2.77±0.17
8.8±0.4
50.84±2.0
0.17±0.01
1.4±0.1
9.2±0.5
2413±49
4966±126
11014±321
ESNN-Ortho
2.56±0.15
8.5±0.4
55.42±2.2
0.13±0.01
1.1±0.1
8.1±0.4
1459±27
2807±66
8821±259
Table 3: Mesh-based physical dynamics. RMSE ( ×10−3 ) for one-step, 50-step, and full autoregressive rollouts on CylinderFlow , DeformingPlate , and Airfoil . Results are mean ± standard deviation over five matched runs. Best values are bold; Diverged denotes an unstable autoregressive rollout. Lower is better.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Spatial–channel factorization of ESNN transport. Each term acts on Vj∈Rn×cv along two independent axes: Sij(k) acts on the ambient spatial dimension by left multiplication, while Mij(k) acts on the vector channels by right multiplication. Summing the separable terms gives the canonical transport Ti←j . ESNN enforces equivariance by constraining the spatial operators to be O(n) -covariant and the channel maps to be invariant.
Figure 5: Charged-particle dynamics without and with a uniform gravitational field. Translation symmetry is preserved, while the ideal unclipped gravity–Coulomb dynamics retain only orthogonal transformations that preserve gtrue .
Figure 6: ModelNet40 graph construction and rotation protocols. Each centered point cloud is converted into a local k -nearest-neighbor graph and evaluated under the z/z , z/SO(3) , and SO(3)/SO(3) train/test regimes.
Figure 7: Mesh-based dynamics benchmarks: CylinderFlow , DeformingPlate , and Airfoil . Models are trained on one-step targets and evaluated autoregressively.
Method
α
Δϵ
ϵHOMO
ϵLUMO
μ
Cv
G
H
⟨R2⟩
U
U0
ZPVE
Units
ma03
meV
meV
meV
mD
mcalmol−1K−1
meV
meV
ma02
meV
meV
meV
Invariant models
Cormorant
85
61
34
38
38
26
20
21
961
21
22
2.03
NMP
92
69
43
38
30
40
19
17
180
20
20
1.50
DimeNet++ †
44
32.6
24.6
19.5
29.7
23
7.56
6.53
331
6.28
6.32
1.21
ComENet †
45
32.4
23.1
19.8
24.5
22
7.98
6.86
259
6.82
6.69
1.20
Appendix
Table 4: QM9 molecular-property prediction. Mean absolute error (MAE) on twelve invariant molecular properties. Baseline values and partition annotations follow Aykent and Xia (2025) ; † marks a different data partition. ESNN results are reported for the validation-selected model for each property, with the best ESNN result shown in bold . Units are given in the second header row. Lower is better.
Task
Model
Transport / training regime
C
L
V
Learning rate
Epochs
Charged N-body
ESNN
Identity
64
6
0
3.38×10−3
10000
ESNN
Diagonal
32
5
1
3.25×10−3
3000
ESNN
Orthogonal
32
7
1
1.92×10−3
3000
ESNN
Radial–Tangential
32
5
1
3.41×10−3
3000
Gravity N-body
ESNN
Diagonal
64
7
1
1.01×10−3
3000
ESNN
Orthogonal
32
7
1
1.24×10−3
3000
Appendix
Table 5: Selected N-body, gravity, and ModelNet40 configurations. C , L , and V denote scalar width, message-passing layers, and vector groups. For ModelNet40, z/z and z/SO(3) use the same z -trained checkpoint.
Symmetry is everywhere in nature and society. Geometric deep learning exploits symmetries in data to improve the performance and efficiency of deep learning systems. In this paper, we extend geometric deep learning to utilize richer symmetry structures. Specifically, we develop order-equivariant neural networks (OENN), which generalize standard graph message passing and sheaf neural networks via the theory of equivariant bundles over face posets (face categories). We (i) characterize all linear order-equivariant maps, (ii) build OENN layers, and (iii) prove universal approximation theorems (UATs) for continuous order-equivariant maps, which are new results even when restricted to sheaf neural networks (for which no UAT was known before). We illustrate the framework on graph and sheaf models. Our results can also be seen as extending the known UAT for graph neural networks to a more general setting that subsumes sheaf neural networks as well. In addition, we show that OENN can be extended further to CENN, Category-Equivariant Neural Network, which gives the general form of equivariant neural networks as well as of equivariant universal approximation theorems, allowing us to leverage categorical symmetry in data (e.g., non-invertible symmetries on multiple objects with compositional relations on those symmetries).
Yoshihiro Maruyama
Department of Mathematical and Information Sciences, Kyoto University, Kyoto, Japan.
We introduce the Clifford Sheaf Neural Network (CSNN), an equivariant sheaf neural network for geometric graphs that places a Clifford algebra on each stalk of a cellular sheaf and transports multivector features along edges. The canonical choice of restriction map for sheaves with algebra-valued stalks is algebra homomorphism. Adding the constraint of equivariance, the naive choice becomes versor conjugation. However, versor conjugation is expressively weak, so we drop algebra homomorphism and arrive at the K-term sandwich. The resulting sheaf Laplacian is positive semidefinite by construction, needs no versor constraint, and still mixes grades. Our main contribution characterizes the resulting family of restriction maps along three axes: which grades a map couples, how much of the endomorphism space it reaches, and how well it is conditioned. The K-term sandwich spans half of the endomorphism space, and in Cl(3, 0, 0) it corresponds to the maps that commute with the central pseudoscalar. The number of terms controls expressivity. CSNN is the reversion member, a first-order model by construction and the grade-mixing corner of this family, developed as a sheaf construction for graph-level equivariant regression.
Graph Neural Networks (GNNs) have become the de facto standard for learning on relational data. While traditional GNNs' message passing is well suited for vector-valued node features, there are cases in which node features are better represented by probability distributions than real vectors. Concretely, when node features are Gaussians, characterized by a mean and a covariance matrix, naively concatenating their parameters into a single vector and applying standard message passing discards the geometric and algebraic structure that governs means and covariances. We propose Gaussian Sheaf Neural Networks (GSNNs), a principled framework that incorporates these inductive biases into graph-based learning. Building on the theory of cellular sheaves, we derive a new Laplacian operator that generalizes the sheaf Laplacian to this setting and preserves its key properties. We complement our theoretical contributions with experiments on synthetic and real-world data that illustrate the practical relevance of GSNNs.
André Ribeiro, Ana Luiza Tenório, Tiago da Silva +1