Temporal graph learning is crucial for dynamic networks where nodes and edges evolve over time and new nodes continuously join the system. Inductive representation learning in such settings faces two major challenges: effectively representing unseen nodes and mitigating noisy or redundant graph information. We propose GTGIB, a versatile framework that integrates Graph Structure Learning (GSL) with Temporal Graph Information Bottleneck (TGIB). We design a novel two-step GSL-based structural enhancer to enrich and optimize node neighborhoods and demonstrate its effectiveness and efficiency through theoretical proofs and experiments. The TGIB refines the optimized graph by extending the information bottleneck principle to temporal graphs, regularizing both edges and features based on our derived tractable TGIB objective function via variational approximation, enabling stable and efficient optimization. GTGIB-based models are evaluated to predict links on four real-world datasets; they outperform existing methods in all datasets under the inductive setting, with significant and consistent improvement in the transductive setting.
Figures & tables
Setting
Model
UCI
Social Evolution
MOOC
Wikipedia
Transductive
JODIE
86.73 ± 1.0
76.74 ± 1.2
79.98 ± 0.4
94.62 ± 0.5
DyRep
54.60 ± 3.1
66.02 ± 0.7
80.45 ± 0.5
94.59 ± 0.2
TGAT
77.51 ± 0.7
63.36 ± 0.2
69.75 ± 0.2
95.34 ± 0.1
GraphMixer
93.40 ± 0.6
89.11 ± 0.2
82.73 ± 0.2
96.89 ± 0.1
GraphMixer+TGSL
88.94 ± 0.9
89.30 ± 0.2
79.59 ± 0.7
97.19 ± 0.4
DyGFormer
95.76 ± 0.2
90.75 ± 0.2
87.23 ± 0.5
98.82 ± 0.1
Table 1: TLP performance (average precision, mean ± std). Best results are marked in bold, second best results are underlined.
Model
Transductive AP
Δ mean
Inductive AP
Δ mean
TGN
80.40 ± 1.4
–
76.70 ± 0.9
–
GTGIB-TGN (rand only)
83.63 ± 1.1
+3.23
79.29 ± 1.0
+2.59
GTGIB-TGN (hop only)
85.08 ± 0.8
+4.68
79.80 ± 0.5
+3.10
GTGIB-TGN ( w/o. enhancer)
82.94 ± 0.5
+2.54
74.47 ± 0.4
-2.23
GTGIB-TGN ( w/o. TGIB)
85.88 ± 0.6
+5.48
79.78 ± 0.5
+3.08
GTGIB-TGN
87.04 ± 0.5
+6.64
81.54 ± 0.5
+4.84
Table 2: Ablation study of GTGIB-TGN on UCI. “Rand only” means only the random structure enhancer is used; “hop only” means only the hop-based structure enhancer is used; “ w/o enhancer” means no sampling is used. “ w/o TGIB” means both random and hop-based sampling but no TGIB is used. “ Δ mean” represents the difference between the mean values of each model and TGN.
#samples
Transductive AP
Δ
Inductive AP
Δ
0
82.94
–
74.47
–
15
83.74
0.80
76.11
1.64
30
86.04
2.30
81.51
5.40
45
88.96
2.92
84.56
3.05
60
92.29
3.33
87.36
2.81
75
92.64
0.35
87.67
0.31
Table 3: GTGIB-TGN on UCI. “#sample” is the total number of neighbors sampled by the structure enhancer. Δ over the last line.
Model
Transductive AUC (Yelp)
Inductive AUC (Yelp)
GIB
77.52 ± 0.4
–
DGIB-Bern
76.88 ± 0.2
–
DGIB-Cat
79.53 ± 0.2
–
GTGIB-TGN
90.23 ± 0.1
82.94 ± 0.2
GTGIB-CAW
88.51 ± 0.1
73.83 ± 0.3
Table 4: Comparison of IB-based methods in TLP (AUC, mean ± std)
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Nodes
Edges
Edge Feature
Node Feature
Time
Duration
Avg.
Dimension
Dimension
Granularity
Degree
Wikipedia
9,227
157,474
172
172
Unix timestamp
1 month
34.14
MOOC
7,145
411,749
4
0
Unix timestamp
17 months
115.15
UCI
1,899
59,835
0
0
Unix timestamp
196 days
63.00
Social Evo.
74
2,099,519
0
2
Unix timestamp
8 months
56,743.76
Appendix
Table 5: Summary of Dataset Statistics
Model
Runtime
Δ %
CAW (backbone)
51.3
65.33%
GTGIB-CAW (rand only)
66.2
19.28%
GTGIB-CAW (hop only)
61.5
13.20%
GTGIB-CAW (TGIB only)
56.5
6.73%
GTGIB-CAW (rand & hop)
69.0
22.90%
GTGIB-CAW (all)
77.3
33.64%
Appendix
Table 6: Runtime and relative overhead ( Δ %) of the CAW backbone and various GTGIB-CAW configurations on an NVIDIA A100 GPU on UCI dataset. The Δ % column indicates the fraction of runtime attributable to the extra GTGIB-CAW module.
Dataset
CAW
GTGIB-CAW
Δ %
GraphMixer
GraphMixer+TGSL
Δ %
Wikipedia
290.9
382.9
31.64%
71.0
1407.7
1882.71%
UCI
140.0
210.47
50.34%
11.3
398.0
3422.12%
MOOC
2213.5
3508.00
58.48%
186.7
5194.30
2682.16%
Social Evo.
425.0
691.25
62.65%
37.0
1650.0
4359.46%
Appendix
Table 7: Mean epoch times (in seconds) on four datasets using one NVIDIA V100 GPU. Comparison of the CAW backbone, GTGIB-CAW, and baselines (GraphMixer backbone and GraphMixer+TGSL). Δ % indicates the runtime increase.