Graph generative models increasingly rely on Graph Transformers (GT) to capture complex dependencies among nodes and edges. While deeper architectures should provide greater expressive capacity and a broader receptive field, their effectiveness can decline with depth: repeated self-attention progressively contracts node representations, impeding information flow and gradient propagation. We analyse this phenomenon from a dynamical systems perspective, focusing on how the denoiser's spectral dynamics affect graph generation. We show that standard GT denoisers become increasingly dissipative as depth grows, leading to vanishing gradients and representation collapse. To isolate the effect of these dynamics, we construct a permutation-equivariant GT with inherently stable, non-dissipative transport. We also introduce a damping mechanism that continuously interpolates between non-dissipative and increasingly contractive regimes, enabling a direct assessment of how dissipation influences generation. Experiments on synthetic and molecular graph generation benchmarks show that the gap between these regimes widens with depth: non-dissipative dynamics preserve representation diversity and gradient flow, sustaining strong generative performance, whereas greater contraction progressively impairs it. These findings identify the denoiser's dynamical regime as a key design factor for deep graph generative models.
Figures & tables
Figure 1: Illustration of node-state trajectories across depth L . (a) Under a standard Graph Transformer denoiser, node states spiral into a single point and node-specific information is lost. (b) Under non-dissipative orthogonal transport (ours), node states rotate while keeping their norms and remain distinct.
Figure 2: Jacobian spectra of the GT by Vignac et al. (2023) for different number of layers L .
Figure 3: Spectral behavior of SGT. (a) Jacobian eigenvalue spectra of SGT with L=8 for different values of γ . Increasing γ progressively moves the spectrum from the unit circle into increasingly contractive regimes. (b) Jacobian singular values at L=32 for SGT and GT proposed by Qin et al. (2025) .
Figure 4: Graphs generated by SGT on Planar (top) and SBM (bottom)
Planar
SBM
Model
V.U.N. ↑
Ratio ↓
V.U.N. ↑
Ratio ↓
Train set
100.0
1.0
85.9
1.0
GraphRNN
0.0
490.2
5.0
14.7
GRAN
0.0
2.0
25.0
9.7
SPECTRE
25.0
3.0
52.5
2.2
EDGE
0.0
431.4
0.0
51.4
Table 1: Graph generation performance on Planar and SBM. For our model, we report results at each depth L . Metrics are computed on 40 generated graphs. Higher V.U.N. is better and lower ratio is better. Best in bold, second best underlined.
ZINC
Model
Val. ↑
Unique. ↑
FCD ↓
GruM
98.7
–
2.26
GBD
97.9
–
2.25
CatFlow
99.2
100.0
13.21
GGFlow
99.6
100.0
1.45
DeFoG (retrained L=8 )
99.2
100.0
1.43
Table 2: Molecule generation on ZINC ( 10,000 samples) and MOSES ( 25,000 samples); DeFoG is retrained by us. MOSES FCD is measured against the scaffold-split test set (TestSF), as for all baselines; Test-split FCD is in Table 3 . Best in bold, second best underlined.
Figure 5: Representation dynamics of trained Planar models across depth. Left: Dong residual between node states. Right: effective rank of the node-feature matrix. Each curve ends at its model depth; the L=32 curves are highlighted. Orthogonal transport mitigates the collapse observed in the deep baseline for these checkpoints.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Test
TestSF
Model
Depth
Filters ↑
FCD ↓
SNN ↑
Scaf ↑
FCD ↓
SNN ↑
Scaf ↑
DeFoG
L=8
99.21
0.818
0.584
85.28
1.339
0.552
12.13
L=16
99.21
1.129
0.591
86.52
1.676
0.556
11.10
SGT
L=8
99.02
0.793
0.587
85.38
1.229
0.556
10.26
L=16
99.21
0.777
0.600
86.67
1.151
0.567
11.09
L=32
99.21
0.705
0.597
90.63
1.101
0.565
11.45
Appendix
Table 3: MOSES benchmark metrics (molsets), test phase on 25,000 generated molecules, 500 sampling steps. FCD, SNN and Scaf are measured against the MOSES test set (Test) and the scaffold test set (TestSF); Filters and Scaf are percentages. Train data scores n random training molecules and is the reference for a perfect model at the same n (Scaf/TestSF is 0 by construction: the scaffold split shares no scaffolds with training). Best model value in bold, second best underlined.
Model
Depth
FCD ↓
SNN ↑
Scaf ↑
Frag ↑
IntDiv ↑
DeFoG
L=8
1.43
0.433
62.17
0.995
0.862
L=16
1.30
0.441
61.50
0.990
0.863
SGT
L=8
0.94
0.434
60.38
0.996
0.863
L=16
0.86
0.442
66.37
0.998
0.864
L=32
0.85
0.452
66.13
0.998
0.868
Train data
0.21
0.483
63.22
1.000
0.869
Appendix
Table 4: ZINC250k, MOSES benchmark metrics (molsets) against the ZINC250k test set, test phase on n=104 generated molecules, single fold. Scaf is a percentage. Train data scores 104 random training molecules, the reference for a perfect model at the same n ; Scaf and IntDiv can exceed it. ZINC250k has no scaffold split, so only the Test variants exist. Best model value in bold, second best underlined.
Component
Specification
Compute node
Platform
Dell PowerEdge XE9640 (BIOS 2.11.2)
CPU
2 × Intel Xeon Platinum 8452Y
36 cores / 72 threads each (144 threads total)
Clock
0.8–3.2 GHz
Cache
3.4 MiB L1d, 2.3 MiB L1i, 144 MiB L2, 135 MiB L3
Appendix
Table 5: Hardware and software configuration of the experimental platform.
Parameters (M)
Time / epoch (s)
Dataset
L
Baseline
SGT
Δ
Baseline
SGT
Planar
8
7.14
6.02
−15.6%
0.4
0.4
16
14.18
11.95
−15.7%
0.8
0.7
32
28.25
23.79
−15.8%
1.5
1.5
SBM
8
7.14
6.03
−15.6%
2.9
2.6
16
14.18
11.95
−15.7%
8.3
7.2
Appendix
Table 6: Model size and training cost. Δ is the SGT’s parameter count relative to the baseline (i.e., the standard GT used in Qin et al. (2025) ) at the same depth. Time per epoch is the median wall-clock of the training-only epochs (no validation or sampling) logged during each campaign run, excluding the first epoch of every process; times are measured on a single GPU of the platform described in Table 5 .
γ0
non-dissipative ⟶ dissipative
0
0.5
1
4
Planar
100.0
97.5
97.5
95.0
SBM
92.5
90.0
90.0
90.0
Appendix
Table 7: V.U.N. (%) of SGT at L=32 with damping γ0 . γ0=0 is exactly orthogonal (i.e., non-dissipative dynamics), while increasing γ0 induces progressively stronger dissipation. 40 generated graphs.
Discrete graph generation has emerged as a powerful paradigm for modeling graph-structured data, yet state of the art models often rely on Graph Transformers or higher order architectures. We revisit this design assumption by introducing GenGNN, a modular message passing backbone for graph generation. GenGNN enables powerful generation by persisting edge fields through latent refinement of coupled node edge graph states, all without requiring global attention. Diffusion models integrating GenGNN achieve over 90 percent validity on standard benchmark datasets, performing within margins of Graph Transformer backbones and achieving up to 2x or even 5x faster inference. Systematic ablations isolate how GenGNN is resilient to oversmoothing during generative denoising, indicating each GenGNN component is necessary for downstream generation quality. Finally, representation-space analysis suggests GenGNN learns functionally similar representations to more theoretically-expressive architectures; even at deeper layers. As such, GenGNN uplifts local message-passing to challenge prevailing assumptions that performant discrete graph generation requires global attention or higher-order representations. Source Code Available Here
Jay Revolinsky, Harry Shomer, Jiliang Tang
Michigan State University East Lansing, Michigan, USA · University of Texas - Arlington Arlington, Texas, USA
Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbations. While recent graph-aware Schrödinger bridge models incorporate topology information directly into their reference dynamics, it is unclear how perturbations of the graph propagate through these dynamics and affect the resulting generated distributions. In this paper, we analyze the structural stability of graph-aware continuous-time generative models whose drift combines a graph filter with a learned graph neural network. We derive explicit Wasserstein stability bounds that quantify the effect of relative graph perturbations on the generated distributions. Motivated by these bounds, we introduce a principled framework for designing stable graph filters that preserve the smoothing behavior of graph heat diffusion, while boosting structural stability. Experiments on synthetic and fMRI signals show our stable filters enhance structural robustness while matching or exceeding the generative quality of the heat equation baseline.
Denoising graphs is a fundamental problem in graph learning and the core operation of graph diffusion models. Attention-based architectures like graph transformers have recently shown promise in denoising graphs. However, our principled understanding of attention-based graph denoising remains limited, making it unclear whether standard attention is the right mechanism for this task. Here we show that, under a denoising objective, linear attention is suboptimal and can only learn an average spectral denoising filter over the training distribution. This creates a fundamental limitation as graphs often vary spectrally across the distribution. To overcome this limitation, we introduce Spectral Attention, which directly utilizes the input graph spectrum and provably outperforms linear attention by a margin governed by the spectral diversity of the distribution. We then derive Graph Convolutional Attention (GCA), a practical and permutation-equivariant realization of this idea that implements spectral denoising through graph-filtered queries and keys. For stochastic block models, GCA provably matches the idealized Spectral Attention mechanism. We further show that the softmax operation, that follows the attention, provides additional denoising by approximately projecting noisy eigenvectors onto the clean eigenspace. Empirically, replacing linear attention with GCA consistently improves graph denoising and diffusion on synthetic and real datasets, with gains strongly correlated with spectral diversity. In DiGress, GCA matches standard graph-transformer performance without computing expensive structural features, and when combined with the recently proposed PEARL positional encodings, avoids explicit eigendecomposition computations resulting in faster inference without degrading quality. The code can be found here: github.com/shervinkhalafi/graph_conv_att
Shervin Khalafi, Igor Krawczuk, Sergio Rozada +3
University of Pennsylvania · King Juan Carlos University · Stanford University