Temporal knowledge graph forecasting aims to infer future relational facts from the temporal structure of observed events. Existing forecasters mainly summarize history through entity states, relation states, paths, or exact recurrence. These views often miss pair-specific transition evidence, that is, the way prior relations between the query actor and a candidate change the odds of the target relation. We introduce BridgeMem, which estimates this quantity as a residual added to the log scores of a frozen full-vocabulary forecaster. For each candidate, BridgeMem retrieves the pair's events that strictly precede t, encodes their relations, directions, and lags, and converts them into a likelihood-ratio correction. A support-adaptive empirical-Bayes reader trusts exact transition counts where they are abundant and backs off to a learned attention estimator where they are sparse. The backbone's own uncertainty gates the correction, so confident queries and candidates without dyadic history are left unchanged. On five benchmarks, BridgeMem improves on the strongest of nine baselines from 2021--2026 in all 20 filtered MRR and Hits@{1,3,10} comparisons, with MRR gains of 0.0213, 0.0164, 0.0216, 0.0112, and 0.0028 over the best prior result. These results show the value of explicit dyadic transition modeling.
Figures & tables
Method
ICEWS 2014
ICEWS 2018
ICEWS 05–15
GDELT
WIKI
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
CyGNet (2021)
.3863
.2871
.4324
.5815
.2767
.1799
.3145
.4702
.4056
.2993
.4576
.6116
.2056
.1280
.2198
.3583
.6580
.5706
.7124
.8251
RE–GCN (2021)
.4221
.3185
.4718
.6215
.3249
.2235
.3664
.5238
.4670
.3618
.5236
.6691
.1995
.1265
.2124
.3422
.7763
.7386
.8034
.8368
TiRGN (2022)
.4478
.3411
.5062
.6493
.3366
.2311
.3811
.5434
.4950
.3861
.5559
.7016
.2173
.1364
.2341
.3778
.8164
.7780
.8509
.8710
CENET (2023)
.5683
.5206
.5857
.6617
.5253
.4831
.5384
.6055
.6914
.6560
.7038
.7595
.6428
.6161
.6491
.6899
.8681
.8660
.8702
.8710
HRI (2024)
.3760
.2970
.4154
.5242
.2839
.2035
.3203
.4378
.4434
.3504
.4963
.6170
.2465
.1667
.2714
.4001
.8155
.7736
.8571
.8701
Table 1: Complete filtered entity-forecasting matrix. Each dataset repeats MRR, H@1, H@3, and H@10. The last block reports BridgeMem and its per-metric improvement over the strongest prior method; all deltas are positive.
Dataset
Params.
Epoch
Trans.
Select+fit
Final refit
ICEWS 2014
83,021
20
88,988
86.2
55.0
ICEWS 2018
91,393
20
448,044
329.7
194.3
ICEWS 05–15
89,783
20
476,338
350.4
217.6
GDELT
86,241
18
2,000,000
1424.2
656.7
WIKI
16,689
15
1,048,280
613.3
340.6
Table 2: Current BridgeMem fitting cost on one NVIDIA H20. All times are in seconds. Added parameters exclude the frozen CENET backbone. Select+fit includes chronological epoch selection and the train-only refit; Final refit trains on train plus validation for the selected epoch budget.
Temporal knowledge graphs (TKGs) represent evolving relational systems, whose underlying data-generating processes often change over time. Yet, TKG forecasting models are commonly evaluated only on empirical benchmark datasets that provide limited insight into the models' robustness to such distribution shifts. Recognising this issue, we study TKG forecasting under controlled shift environments using a synthetic TKG generator that encodes three temporal and structural properties -- recurrence, homophily, and periodicity -- as data-generating mechanisms. This allows us to evaluate seven forecasting architectures under stationary and shifting regimes. Our experiments suggest that robustness in TKG forecasting is highly signal-dependent. Recurrence-based and periodic regularities are largely recoverable under stationary conditions, and simple memory-based baselines can be competitive when recurrence dominates the data. However, structural breaks reveal limitations in model adaptivity, with shifts in latent entity-community structure posing the strongest challenge in our study. Overall, our findings improve the understanding of the capabilities and limitations of current TKG models confronted with temporal distribution shifts.
Konrad Özdemir, Julia Gastinger, Lukas Kirchdorfer +1
Data and Web Science Group, University of Mannheim, Germany · SAP Signavio, Walldorf, Germany
Aggregate performance on continuous-time dynamic graphs (CTDGs) combines, in a single score, the portion attributable to known temporal regularities and the additional predictive power of neural models. This study separates the two at the query level. We construct a mechanism-constrained predictor that uses pair recurrence, recency and history position, renewal patterns, and short sequential transitions while learning the compatibility within each mechanism. Across four CTDG datasets, this predictor recovers a substantial portion of the performance of strong neural baselines, and the recovered performance quickly saturates with a small, dataset-specific set of explicit mechanisms. Neural residuals concentrate on queries for which the positive and negative candidates have similar mechanism-execution profiles. Allowing conditional interactions among mechanisms is more effective than simply reweighting their existing contributions. Conditioning the contribution of one mechanism on the execution state of another recovers 54.9-73.2% of the original neural-only queries and improves overall paired accuracy on all four datasets. Although the magnitude of the effect varies across datasets, these results show that the performance gap of neural CTDG models need not be treated solely as an opaque difference in representational capacity. At least part of the gap is localized to queries with similar candidate execution profiles and can be functionally explained by conditional coordination among known, low-dimensional mechanisms.
Minwoo Yu, Young-guk Ha
Smart Computing Laboratory, Department of Computer Science & Engineering, Konkuk University, Seoul 05029, Republic of Korea
Temporal link prediction on the Temporal Graph Benchmark 2.0 (TGB 2.0) faces a scalability ceiling: on the benchmark's three largest datasets, every existing embedding method runs out of memory or exceeds the time budget. These large-scale graphs are the ones nearest real deployment scale, so failing on them is a real production limitation. EdgeReMIND sets the highest reported test mean reciprocal rank (MRR) on six of eight TGB 2.0 datasets and is the only relation-aware method that runs on all of them. This linear memorization model, with learned per-relation weights over data-calibrated features, is therefore not merely a fallback where embeddings fail but a practical state-of-the-art baseline across the benchmark.