Graph neural networks (GNNs) are often called long-range because their architecture can connect distant nodes, but this does not show whether they use distant information correctly. We introduce a framework that measures how strongly inputs at each graph distance affect predictions and separates limitations due to architecture, finite approximation, training, and numerical execution. Our analysis shows that local message-passing can spread influence slowly, so a finite implementation may rely mainly on nearby inputs even when the ideal computation uses the whole graph. We also explain why mathematically equivalent filters can differ in how easily they are learned and how reliably they run. Across controlled tasks, models with similar architectural reach use distant information very differently, while low average error can hide failures on distant interactions. Together, these results show that long-range capability depends on learning to use information at the distances required by the task and preserving that use during computation.
Figures & tables
Figure 1: Architectural support determines which node pairs may interact. When an ideal operator exists, a finite implementation approximates it; training may fail to learn the profile required by the task, and approximations and rounding errors during inference may further alter the learned profile.
Figure 2: Left : Marked PathCopy accuracy; shading marks training distances and the dashed line marks the support limit at 20 hops. Middle : aggregate error on ResolventCopy hides a tail that finite models miss completely. Right : in RingTransfer, a rational fit matches the frequencies present in the training graph but becomes inaccurate at frequencies that graph does not contain as conditioning worsens.
Figure 3: Top-left : copying accuracy on the test set. Lines average five runs; shading marks the trained distances 2–10. The horizontal line marks five-class chance and the vertical line marks the local models’ 20-hop support. Top-right : RangeProfileCopy test NMSE. Bottom : differences between each model and its matched control in prediction error and realized range R90 on the long-range resolvent. Open markers denote matched controls, filled blue markers denote models with the mechanism being tested, and error bars show sample standard deviations across three runs.
Figure 4: Left : mean test error across runs; lower is better. Right : how often each model correctly identifies which atom in a matched pair has the larger charge when it sees the complete molecules (blue) or only the matched local views (orange). A blue–orange gap means that the rest of the molecule helps the model. The dashed line marks the 0.5 local-only baseline.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Representative graphs from each task. (a) The distance-20 RingTransfer cycle, with the antipodal source and query distinguished from ordinary nodes. (b) A representative Marked PathCopy sample showing the source, query, three distractors, and padding nodes at their sampled positions. (c) The path and ladder topologies used for RangeProfileCopy; node color shows the observed first input channel. (d) Two molecules from the validation split of the ECHO-Charge dataset. Element color identifies atomic number, bond width identifies bond type, and the layout uses observed bond lengths; every atom is a charge-prediction target.
Task or test
Main question
RingTransfer
How does a filter’s mathematical form affect fitting and numerical accuracy?
Marked PathCopy
Can the model copy a source label from increasingly distant nodes?
ResolventCopy
Does the model recover weak distant effects that average error can hide?
RangeProfileCopy
Does the model weight nearby and distant inputs as the task requires?
ECHO-Charge-noCOM
How accurately does the model predict atomic charges in real molecules?
RemotePairs
Does the rest of the molecule improve predictions when local surroundings are identical?
Appendix
Table 1: Tasks and the questions they address.
Model
Support or order
Propagation
Parameters
Chebyshev
degree 20
polynomial
87,237/86,660
Monomial
degree 20
polynomial
87,237/86,660
Finite ARMA
20 stacks, 20 steps
tied recurrence
88,089/87,746
Exact bank
8 fixed poles
equilibrium solves
90,095/89,203
20-step restart
8 poles, 20 steps
matched recurrence
90,095/89,203
Appendix
Table 2: Controlled filter models. Parameter counts are shown as PathCopy/ResolventCopy and include the task-specific encoder and decoder.
PathCopy accuracy
ResolventCopy NMSE
Model
d=10
d=15
d=20
α=.95
α=.995
Tail error
Exact bank
1.000
1.000
1.000
4.3×10−7
1.8×10−6
2.03×10−5
Chebyshev
1.000
1.000
0.997
2.6×10−6
2.75×10−2
1.000
Finite ARMA
1.000
0.247
0.195
3.1×10−4
1.22×10−1
1.000
20-step restart
1.000
0.201
0.210
4.3×10−3
2.91×10−1
1.000
Monomial
0.995
0.193
0.193
5.1×10−4
1.54×10−1
1.000
Appendix
Table 3: Controlled-filter results, averaged across three independent training runs. PathCopy models train at distances 2–10. ResolventCopy NMSE is evaluated on 120-node paths, and tail error is measured strictly beyond hop 20 for α=0.995 .
Figure 6: Fixed-pole models with their learned parameters held fixed. Blue uses the Exact-bank parameters and orange the 20-step-restart parameters; the models were trained separately. “32-bit equilibrium” runs both sets to equilibrium, including the orange one. Numbered points run the same parameters for the indicated number of 64-bit restart steps. “Original model” uses 32-bit equilibrium for blue and 20 32-bit steps for orange. Relative difference is measured from the same model’s 64-bit equilibrium; lines and bands show the mean and sample standard deviation over three runs.
α
Evaluation
Rel. difference
NMSE
R90
Mass >20
Tail error
.95
64-bit equilibrium
0
5.53×10−7
7
.00183
.191
.95
20 steps
.316
.100
6
0
1.000
.95
160 steps
.00847
7.50×10−5
7
.00120
.0272
.995
64-bit equilibrium
0
7.35×10−7
23
.1259
.00219
.995
20 steps
1.586
2.514
7
0
1.000
.995
160 steps
.448
.200
18
.0559
.658
Appendix
Table 4: Selected values for the Exact-bank models (blue curves in Figure 6 ) on 120-node ResolventCopy paths. Relative difference is measured from the same model’s 64-bit equilibrium. Profile statistics use an endpoint impulse, and tail error is measured beyond hop 20. Entries are means over three runs.
Target
Fit
Synthesis condition
High-precision RMSE
FP32 RMSE
T10
10 real poles
2.74×106
0.986
0.985
T15
15 real poles
5.12×109
2.38
32.8
T20
20 real poles
9.91×1012
6.63
5.24×104
Resolvent α=0.995
degree-20 Chebyshev
1.41
1.95×10−2
1.95×10−2
Resolvent α=0.995
20 real poles
6.76×1012
5.39×10−4
5.56×10−4
Appendix
Table 5: Compact cross-spectrum results on the dense interval for the reported real-pole configurations. The smooth-resolvent rows compare the degree-20 Chebyshev and rational fits.
Model
Primary propagation mechanism
PathCopy params.
RangeProfile params.
GatedGCN
20 residual local steps with persistent directed-edge states
108,709
108,548
Stable-ChebNet
one stabilized degree-20 Chebyshev block
97,493
97,012
AMP
20 GatedGCN steps, global depth posterior, and channelwise message gates
109,991
109,830
GraphGPS
four blocks combining GatedGCN and within-graph global attention
87,085
86,884
MP-SSM
one sequential normalized-adjacency state recurrence with 20 updates
95,461
104,804
MLP
two pointwise layers and no graph communication
98,069
96,988
Appendix
Table 6: Models in the GNN-family comparison. Parameter counts include the common encoder and decoder. GraphGPS+RWSE is a structural-information control.
Task
Quantity
Standard–direct difference
Metric
Direct solve
Standard solver
5 iterations
RangeProfileCopy
Local
.005
NMSE
.0926
.0925
.196
Inverse-distance
3.8×10−5
NMSE
.164
.164
.242
Resolvent
.012
NMSE
.000117
.000260
.240
Resolvent influence
–
R90
20.1
18.8
38.8
Appendix
Table 12: Effect of the Tikhonov solver on RangeProfileCopy. “Direct solve” is the 64-bit reference, the standard solver uses 30 iterations, and “5 iterations” deliberately stops early. Prediction rows report test NMSE and the relative difference between the standard and direct predictions; the final row reports the average resolvent R90 in graph hops. Entries are means over three runs.
Control → full
NMSE control
NMSE full
R90 control
R90 full
Δ cosine
Full M>20
GatedGCN-5 → GatedGCN-20
0.458±0.012
0.126±0.006
4.19±0.01
10.06±0.12
+0.1313
0.000
Vanilla ChebNet → Stable-ChebNet
0.024±0.001
0.023±0.000
13.75±0.07
13.78±0.07
−0.0001
0.000
GatedGCN-20 → AMP
0.126±0.006
0.456±0.016
10.06±0.12
3.65±0.19
−0.1551
0.000
GraphGPS-local → GraphGPS
0.542±0.012
0.265±0.001
3.66±0.09
48.00±0.09
+0.0932
0.385
MP-SSM-local → MP-SSM-20
0.823±0.003
0.160±0.002
1.00±0.00
7.53±0.04
+0.4072
0.000
Appendix
Table 13: Architectural contrasts on hard-resolvent RangeProfileCopy. From top to bottom, the pairs vary GatedGCN depth and width, ChebNet stabilization, fixed versus adaptive depth, GraphGPS global attention, and MP-SSM state recurrence. NMSE is pooled over the test set; influence statistics average 2,000 topology–size–target cases per run. Values are means over three independent runs; entries displaying ± include the sample standard deviation.
Model
Output drift (%)
Δ NMSE
R90 retention (%)
Exact R90 (%)
ΔEtail
GatedGCN-20
3.78±0.11
−0.0012
100.02±0.15
89.5±3.0
+0.000
Stable-ChebNet
17.76±2.69
+0.0768
99.96±0.05
95.9±1.8
+0.000
AMP
4.42±1.53
+0.0001
99.98±0.12
99.5±0.1
+0.000
GraphGPS
1.56±0.10
−0.0001
100.00±0.03
95.4±1.6
−0.037
MP-SSM-20
7.22±0.14
+0.0205
99.81±0.24
98.4±2.2
+0.000
SONAR
5.35±0.34
+0.0037
100.03±0.05
98.0±0.7
+0.000
Appendix
Table 14: BF16 precision replay relative to FP64 on hard-resolvent RangeProfileCopy. Each run replays 256 graphs and 512 influence targets using the same trained parameters. Values are means over three independent runs; entries displaying ± include the sample standard deviation.
Figure 7: Complementary hard-resolvent and AMP diagnostics. (a) BF16 output drift relative to FP64 for different architectures; bars show sample standard deviations across runs. (b) AMP’s learned global depth posterior; lines are run means and bands are sample standard deviations. (c) AMP’s mean message-gate profile; lines average run-level means and bands average the within-run 10th–90th percentiles.
Model
Width
Parameters
Learning rate
GatedGCN-3
176
506,886
10−3
GatedGCN-20
72
542,958
10−3
Stable-ChebNet
208
499,518
10−3
GraphGPS
96
503,526
10−3
MP-SSM-20
320
501,126
3×10−4
A-DGN
424
501,598
10−3
Appendix
Table 15: Model sizes and learning rates for molecular training. Width and total trainable parameter counts include the encoder and prediction head.
Model
Test MAE
Full acc.
Crop acc.
Full − crop
GRIT
0.0059
0.690 [0.651, 0.726]
0.500 [0.500, 0.500]
0.190 [0.151, 0.226]
GatedGCN-20
0.0060
0.624 [0.584, 0.663]
0.500 [0.500, 0.500]
0.124 [0.084, 0.163]
GraphGPS
0.0064
0.639 [0.595, 0.685]
0.500 [0.500, 0.500]
0.139 [0.095, 0.185]
A-DGN
0.0066
0.638 [0.598, 0.678]
0.500 [0.500, 0.500]
0.138 [0.098, 0.178]
GatedGCN-3
0.0069
0.507 [0.497, 0.517]
0.500 [0.500, 0.500]
0.007 [-0.003, 0.017]
SONAR
0.0077
0.615 [0.573, 0.653]
0.500 [0.500, 0.500]
0.115 [0.073, 0.153]
Appendix
Table 16: RemotePairs full-molecule and radius-crop intervention. Intervals are paired hierarchical 95% bootstrap intervals.
Model
Full balanced acc.
Crop balanced acc.
A-DGN
0.640 [0.600, 0.680]
0.500 [0.500, 0.500]
GRIT
0.687 [0.648, 0.724]
0.500 [0.500, 0.500]
GatedGCN-20
0.624 [0.584, 0.663]
0.500 [0.500, 0.500]
GatedGCN-3
0.507 [0.497, 0.517]
0.500 [0.500, 0.500]
GraphGPS
0.634 [0.591, 0.680]
0.500 [0.500, 0.500]
MP-SSM
0.594 [0.555, 0.631]
0.500 [0.496, 0.500]
Appendix
Table 17: RemotePairs balanced charge-order accuracy. Entries are medians with paired hierarchical 95% bootstrap intervals. The predeclared claim rule uses the full-molecule median.
Model
Radius
Pairs
Full acc.
Crop acc.
Full − crop
A-DGN
2
348
0.730
0.500
+0.230
A-DGN
3
207
0.626
0.500
+0.126
A-DGN
4
165
0.558
0.500
+0.058
GRIT
2
348
0.767
0.500
+0.267
GRIT
3
207
0.652
0.500
+0.152
GRIT
4
165
0.648
0.500
+0.148
Appendix
Table 18: RemotePairs intervention by matching radius, pooled across the three seeds.
Department of Computer Science, University of Trento, Italy. · Fondazione Bruno Kessler, Italy. · Department of Computer Science, University of Pisa, Italy.