Do Temporal Link Predictors Need Learned Memory? A Smoothed-Count Baseline with a Handful of Parameters
Authors: Lisi Qarkaxhija, Ingo Scholtes
Organizations: Chair of Machine Learning for Complex Networks Center for Artificial Intelligence and Data Science (CAIDAS) Julius-Maximilians-Universität Würzburg, DE
Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal link predictor based on statistical language modelling. It pools transition and co-occurrence counts across sources to predict links that a source has never formed. We smooth sparse estimates using destination frequencies or Kneser-Ney continuation counts. A shared log-linear rule combines these estimates with popularity, source history, and recency, without node embeddings. In our main evaluation, the model achieves the highest MRR among the compared methods on 7 out of 16 datasets from TGB and TGB-Seq. It also outperforms EdgeBank and Base3 on all 16 datasets and the heuristic family on 14. These gains extend to datasets designed to limit repeated edges. With only 9--13 learned parameters, our model provides a simple and competitive baseline for evaluating future neural temporal link predictors.
Figures & tables
Ours
Best neural
Dataset
EdgeBank
Heuristic
Base3
marginal
continuation
MRR
Method
wiki
64.40
82.06
73.26
82.84 ± 0.04
82.72 ± 0.10
82.70
TPNet
uci
32.40
52.75
35.22
51.57 ± 0.32
51.25 ± 0.26
24.50
TNCN
enron
15.60
84.66
48.65
86.82 ± 0.01
86.83 ± 0.03
37.90
TNCN
subreddit
59.22
74.19
74.25
74.90 ± 0.06
74.79 ± 0.04
69.60
TNCN
lastfm
2.63
16.58
11.37
32.84 ± 0.18
32.87 ± 0.12
15.60
TNCN
Table 1: Test MRR (%) on TGB. Our entries are means and sample standard deviations over five runs. EdgeBank is the stronger of its two variants. Best neural gives the strongest reported neural benchmark score (Appendix A ). Bold marks the highest MRR per row.
Ours
Best neural
Dataset
EdgeBank
Heuristic
Base3
marginal
continuation
MRR
Method
GoogleLocal
1.96
18.39
3.97
31.27 ± 0.02
31.42 ± 0.03
62.88
SGNN-HN
Yelp
9.77
30.81
11.99
53.21 ± 0.06
52.16 ± 0.09
72.69
CRAFT
Taobao
20.28
42.99
20.17
59.58 ± 0.03
58.90 ± 0.06
70.68
CRAFT
ML-20M
1.94
17.71
6.64
39.24 ± 0.15
39.00 ± 0.10
35.91
CRAFT
Flickr
1.96
38.38
34.81
62.68 ± 0.07
62.66 ± 0.08
62.34
CRAFT
Table 2: Test MRR (%) on TGB-Seq. Columns follow Table 1 . Count baselines use the TGB-Seq candidate sets.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Estimator
a
b
c
d
P^(z)
0.20
0.30
0.30
0.20
P^c(z)
0.14
0.43
0.29
0.14
P^(z∣u)
0.20
0.38
0.38
0.03
P^(z∣v1=c)
0.10
0.65
0.15
0.10
P^(z∣v2=b)
0.07
0.10
0.77
0.07
Appendix
Table 3: Smoothed estimates on the worked stream of Figure 1 , computed by the model implementation ( α=1 , counts undecayed). Conditional rows use marginal backoff. Bold marks the candidates each estimator favours.
Dataset
Marginal backoff
Continuation backoff
wiki
all three
all three
uci
all three
all three
enron
all three
all three
subreddit
reference
global recency
lastfm
all three
all three
review
co-occurrence
co-occurrence
Appendix
Table 4: Feature sets used in the main comparisons. Entries name additions to the reference configuration. These choices remain fixed across five runs.
Best neural
Dataset
Updates
EdgeBank
Heuristic
Base3
Ours
MRR
Method
wiki
Event
64.40
82.06
73.26
82.84 ± 0.04
82.70
TPNet
Batch 200
57.10
72.90
68.07
74.08 ± 0.09
82.70
TPNet
uci
Event
32.40
52.75
35.22
51.57 ± 0.32
24.50
TNCN
Batch 200
22.16
37.49
28.22
38.55 ± 0.16
24.50
TNCN
enron
Event
15.60
84.66
48.65
86.83 ± 0.03
37.90
TNCN
Appendix
Table 5: Test MRR (%) on TGB under event-wise and batch-200 updates. EdgeBank, Heuristic, Base3, and Ours share the listed schedule and candidates. Our scores are means and sample standard deviations over five runs. Neural reference scores retain their original protocols and are repeated for comparison. Bold marks the highest MRR in each row.
Best neural
Dataset
Updates
EdgeBank
Heuristic
Base3
Ours
MRR
Method
GoogleLocal
Event
1.96
18.39
3.97
31.42 ± 0.03
62.88
SGNN-HN
Batch 200
1.96
18.39
3.73
29.00 ± 0.01
62.88
SGNN-HN
Yelp
Event
9.77
30.81
11.99
53.21 ± 0.06
72.69
CRAFT
Batch 200
9.76
30.81
11.88
52.26 ± 0.05
72.69
CRAFT
Taobao
Event
20.28
42.99
20.17
59.58 ± 0.03
70.68
CRAFT
Appendix
Table 6: Test MRR (%) on TGB-Seq under event-wise and batch-200 updates. Columns and conventions follow Table 5 . Neural reference scores retain their original protocols.
Dataset
Event-wise
Strict timestamp
Change (points)
TGB
wiki
82.84 ± 0.04
82.84 ± 0.04
0.00 ± 0.00
uci
51.57 ± 0.32
51.43 ± 0.32
-0.14 ± 0.00
enron
86.83 ± 0.03
53.50 ± 0.18
-33.33 ± 0.16
subreddit
74.90 ± 0.06
74.89 ± 0.06
0.00 ± 0.00
lastfm
32.84 ± 0.18
32.78 ± 0.18
-0.06 ± 0.00
Appendix
Table 7: Strict-timestamp sensitivity of our model. Scores are test MRR (%), with means and sample standard deviations over five runs. Changes from event-wise evaluation use unrounded scores paired by seed. Configurations, weights, and candidates are fixed.
Temporal link prediction on the Temporal Graph Benchmark 2.0 (TGB 2.0) faces a scalability ceiling: on the benchmark's three largest datasets, every existing embedding method runs out of memory or exceeds the time budget. These large-scale graphs are the ones nearest real deployment scale, so failing on them is a real production limitation. EdgeReMIND sets the highest reported test mean reciprocal rank (MRR) on six of eight TGB 2.0 datasets and is the only relation-aware method that runs on all of them. This linear memorization model, with learned per-relation weights over data-calibrated features, is therefore not merely a fallback where embeddings fail but a practical state-of-the-art baseline across the benchmark.
Aggregate performance on continuous-time dynamic graphs (CTDGs) combines, in a single score, the portion attributable to known temporal regularities and the additional predictive power of neural models. This study separates the two at the query level. We construct a mechanism-constrained predictor that uses pair recurrence, recency and history position, renewal patterns, and short sequential transitions while learning the compatibility within each mechanism. Across four CTDG datasets, this predictor recovers a substantial portion of the performance of strong neural baselines, and the recovered performance quickly saturates with a small, dataset-specific set of explicit mechanisms. Neural residuals concentrate on queries for which the positive and negative candidates have similar mechanism-execution profiles. Allowing conditional interactions among mechanisms is more effective than simply reweighting their existing contributions. Conditioning the contribution of one mechanism on the execution state of another recovers 54.9-73.2% of the original neural-only queries and improves overall paired accuracy on all four datasets. Although the magnitude of the effect varies across datasets, these results show that the performance gap of neural CTDG models need not be treated solely as an opaque difference in representational capacity. At least part of the gap is localized to queries with similar candidate execution profiles and can be functionally explained by conditional coordination among known, low-dimensional mechanisms.
Minwoo Yu, Young-guk Ha
Smart Computing Laboratory, Department of Computer Science & Engineering, Konkuk University, Seoul 05029, Republic of Korea
Temporal link prediction (TLP) is typically evaluated by predictive performance on unseen edges, but this criterion can conflate predictive accuracy with recovery of the underlying causal mechanism. In stochastic models, Fisher information governs the Cramér--Rao (CR) bound on parameter estimation error: higher Fisher information permits more accurate parameter recovery. We show that, under comonotonicity conditions between Fisher information and entropy, binary logistic models exhibit an estimation--prediction tradeoff: regimes with higher Fisher information, and hence smaller CR bounds, also have higher irreducible predictive entropy. To study this tradeoff in TLP, we introduce a probabilistic causal generator for temporal graphs with transient edges and known ground-truth causal structure, and validate the phenomenon empirically.
Aniq Ur Rahman
Department of Engineering Science, University of Oxford, Oxford, OX1 3PJ, UK