GeoGAE: Scalable Graph-Level Autoencoding via Hyperball Cloud Representations
Authors: Radosław Nowak, Anna Bielawska, Bogusz Stefańczyk, Maciej Sanocki, Paweł Wawrzyński
Organizations: Institute of Theoretical and Applied Informatics Polish Academy of Sciences · IDEAS Research Institute · Faculty of Mathematics, Informatics and Mechanics Warsaw University of Technology
Embedding structured objects into Euclidean spaces has enabled a wide range of successful machine learning applications. Such objects include words, documents, image patches, time series, and graph nodes. In contrast, embedding entire graphs remains a challenging problem. Existing methods either sustain the original order of the graph nodes or match the output nodes to the input ones, both of which create scalability issues. In this work, we propose a graph representation as a cloud of hyperballs, which allows us to define a specific, typically unique, node ordering. Based on this representation, we propose GeoGAE, an autoencoder, in which the Transformer encoder translates a hyperball cloud into a graph-level embedding, and the Transformer decoder translates the graph-level embedding back into the graph. This formulation enables the model to capture both the global graph structure and local relational patterns. We evaluate our method on multiple graph datasets, spanning various domains. The results demonstrate effectiveness of our method in encoding and reconstructing graphs from their embeddings.
Figures & tables
Dataset
Graphs
Avg. nodes
Avg. edges
MUTAG
188
17.90
19.80
AIDS
2,000
15.69
16.20
IMDB-BIN
1,000
19.77
96.53
QM9
129,433
18.03
18.63
SYNTHETIC NEW
300
100.00
196.25
COLLAB
5,000
74.49
2,457.78
Table 1: Basic statistics of datasets used in our benchmark.
Dataset
tok-d
emb
geom-d
encod
decod
lr
enc-h
dec-h
batch
MUTAG
64
128
4
1024×4
256×2
1e-3
8
1
2
AIDS
64
192
6
1024×4
512×2
1e-3
8
1
32
IMDB-BIN
64
192
6
512×8
1024×6
1e-3
2
1
32
QM9
64
192
4
1024×6
1024×4
1e-3
4
4
256
SYNTH. NEW
64
192
6
2048×8
1024×2
5e-4
8
4
8
COLLAB
96
576
9
256×8
2048×4
1e-3
4
1
32
Table 2: Hyperparameters of GeoGAE: tok-d – token dimension, emb – graph embedding size (integer multiple of tok-d), geom-d – dimension of the geometric embedding, encod – width and depth (width × number of layers) of the encoder hidden layers, decod – width and depth (width × number of layers) of the decoder hidden layers, lr – learning rate, enc-h – number of attention heads in the encoder, dec-h – number of attention heads in the decoder, batch – batch size. All datasets except REDDIT-BINARY use a bundle size of 1; REDDIT-BINARY uses a bundle size of 4.
Dataset \ method
PIGVAE [%]
ReGAE [%]
GRALE [%]
GeoGAE [%]
MUTAG
43.5
±
14.9
60.1
±
3.5
6.9
±
3.8
61.3
±
5.4
AIDS
60.2
±
4.9
83.6
±
2.6
1.2
±
0.6
84.1
±
1.7
IMDB-BIN
87.3
±
2.0
90.0
±
2.6
54.8
±
7.7
94.0
±
1.0
QM9
0.0
±
0.0
99.9
±
0.0
98.6
±
0.5
99.4
±
0.2
SYNTHETIC NEW
0.4
±
0.6
16.1
±
3.0
7.3
±
0.7
50.1
±
0.4
COLLAB
52.2
±
5.0
78.0
±
2.4†
OOM
90.9
±
0.7
Table 3: Comparison of methods: F1 score on the test set. After ± we put the standard deviation. †Score from 3 seeds, other 2 got NaN loss. ‡Score taken as reported in the original paper, due to training instability.
Dataset
Mean size error
Dataset
Mean size error
MUTAG
0.12±
0.08
AIDS
0.23±
0.09
IMDB-BIN
0.27±
0.05
QM9
0.01±
0.00
SYNTHETIC NEW
0.00±
0.00
COLLAB
1.45±
0.34
REDDIT-BIN
16.25±
1.50
Table 5: GeoGAE: Mean size error — the average absolute size difference between the target and predicted graphs over the average target graph size. Calculated for experiments for pure graph topology deconstruction, without features encoding and decoding enabled.
Movie collaboration ego-networks. Nodes are actors; an edge connects actors appearing in the same movie. Graphs come from Action and Romance movies.
Graph classification
2 classes
No native node labels, node attributes, edge labels, or edge attributes
QM9
Small organic molecules. Nodes are atoms and edges are chemical bonds; optimized 3-D coordinates and quantum-chemical properties are supplied. Molecules contain H, C, N, O and F, with at most nine non-hydrogen atoms.
Graph-level regression
Multiple continuous targets
Atom types/features, bond types and 3-D coordinates
SYNTHETIC NEW
Artificial graphs designed for controlled graph-classification experiments. Each graph has exactly 100 vertices and approximately 196 edges.
Graph classification
2 classes
Discrete node labels and one-dimensional continuous node attributes; no edge labels
COLLAB
Scientific collaboration ego-networks. Nodes are researchers and edges represent co-authorship. Graph labels correspond to three physics research fields.
Graph classification
3 classes
No native node or edge features/labels
Appendix
Table 6: Datasets
Cluster name
GPU device
GPU driver version
CPU device
Operating system
Total RAM
Cluster 1
NVIDIA A100 PCIe 80GB
565.57.01
AMD EPYC 7713 64-Core Processor
Linux-6.8.0-64-generic-x86_64-with-glibc2.39
2048 GB
Cluster 2
DGXA100 920-23687-2531-001
580.173.02
AMD EPYC 7742 64-Core Processor
Linux-5.15.0-1107-nvidia-with-glibc2.34
1024 GB
Cluster 3
AMD Radeon RX 9070 XT (Navi 48)
Mesa 26.2.2 / ROCm 7.8.0
AMD Ryzen 9 9950X3D 16-Core Processor
Linux-6.18.49-1-MANJARO-x86_64-with-glibc2.44
64 GB
Cluster 4
NVIDIA RTX PRO 500 Blackwell Generation
580.173.02
Intel(R) Core(TM) Ultra 5 235H 14-Core Processor
Linux-7.0.0-28-generic-x86_64-with-glibc2.39
30 GB
Appendix
Table 7: Specification of the computational clusters used in the experiments.
Dataset
No features
With feature decoding
MUTAG
≈24 min
≈25 min
AIDS
≈2h 55 min
≈18h 4 min
IMDB-BIN
≈1h 50 min
N/A
SYNTHETIC NEW
≈44 min
N/A
COLLAB
≈23h 37 min
N/A
QM9
≈29h 16 min
≈38h 16 min
Appendix
Table 8: Runtime comparison for GeoGAE across datasets with and without feature decoding, calculated on Cluster 1 . We report results averaged over 5 seeds. N/A denotes where computations were not applicable (datasets with no node or edge features).
Dataset
emb
hid
ppf-h
heads
layers
k-ls
p-ls
lr
w-decay
MUTAG
64
128
512
4
4
1e-3
1e-1
1e-4
0.0
AIDS
64
128
512
4
4
1e-3
1e-1
1e-4
0.0
IMBD-BIN
16
256
512
4
4
1e-2
1e-1
1e-4
1e-4
QM9
32
128
256
2
1
1e-3
5e-1
1e-4
0.0
SYNTHETIC
16
256
256
8
4
1e-2
1.0
1e-4
0.0
COLLAB
32
128
256
2
1
1e-3
5e-1
1e-4
0.0
Appendix
Table 9: Hyperparameters of PIGVAE: emb – graph embedding size, hid – hidden layer size, ppf-h – transformer layer size, heads – number of transformer heads, layers – number of transformer encoder/decoder layers, k-ls – kld loss scale, p-ls – permutation loss scale, lr – learning rate, w-decay – weight decay.
Dataset
emb
encod
decod
block
l-r
w-decay
MUTAG
160
2048
2048
4
3e-4
1e-3
AIDS
160
2048
2048
4
3e-4
1e-3
IMBD-BIN
160
2048
4096
4
5e-4
1e-4
QM9
160
1024
1024
32
3e-4
1e-3
SYNTHETIC
160
2048
4096
8
1e-3
1e-3
COLLAB
604
2048, 1536
4096
16
3e-4
1e-3
Appendix
Table 10: Hyperparameters of ReGAE: emb – embedding size, encod – sizes of encoder hidden layers, decod – sizes of decoder hidden layers, block – block size, l-r – learning rate, w-decay – weight decay of optimizer. For all experiments we set 0.5 as the mask weight and 0.2 as the embedding norm weight for the loss calculation. We use ELU ( Clevert et al., 2016 ) as the activation function.
Dataset
Layers
Heads
Node Dim
Edge Dim
Latent
Matcher
LR
Dropout
Batch
MUTAG
3
8
64
×
64
null
×
64
256
Soft
1e-4
0.0
64
AIDS
7
8
128
×
128
128
×
128
128
Soft
2e-5
0.1
8
IMDB-BIN
3
4
64
×
64
32
×
32
64
Sink
2e-5
0.0
4
QM9
6
8
64
×
128
64
×
64
128
Sink
1e-4
0.0
256
SYNTH. NEW
5
8
64
×
64
32
×
32
64
Sink
1e-4
0.0
8
COLLAB
3
4
64
×
128
64
×
64
64
Sink
1e-4
0.0
1
Appendix
Table 11: Hyperparameters of GRALE: Layers – number of Evoformer layers, Heads – attention heads, Node/Edge Dim – hidden × model dimension, Latent – graph embedding dimension ( dg ), Matcher – node matching operator (Sink: Sinkhorn, Soft: SoftSort), LR – learning rate, Dropout – attention/MLP dropout, Batch – batch size.
Configuration
Dataset
F1 [%]
Avg. size diff
Baseline results
MUTAG
61.29±5.39
0.12±0.08
AIDS
84.12±1.67
0.23±0.09
QM9
99.38±0.18
0.00±0.00
Using VAE
MUTAG
60.78±2.01
0.24±0.12
AIDS
82.17±1.98
0.61±0.16
QM9
95.88±4.76
0.00±0.00
Appendix
Table 12: The table presents results of ablation study experiments in comparison with the base GeoGAE results.