HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
Authors: Yansong Liu, Rui Liu, Yuan Zuo, Hongwei Zhao, Da Fu, Fuwei Zhang, Fuzhen Zhuang, Yong Chen, +1 more
Organizations: Beihang University Beijing, China · Beijing University of Posts and Telecommunications Beijing, China · Hubei Engineering University Xiaogan, Hubei, China
Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over long horizons. We address this mismatch by formulating extrapolative TKGR as continual learning over streaming snapshots. Under this view, effective extrapolation must jointly handle current dynamics, stable knowledge, and recurring historical evidence. Based on these requirements, we propose History-enhanced Two-Step Continual Learning (HiTS-CL), a backbone-agnostic continual learning framework for extrapolative TKGR. HiTS-CL tracks current dynamics via continual fine-tuning, preserves stable knowledge via multi-teacher adaptive distillation, and retains recurring historical evidence via a selective memory of recent and frequent facts. We integrate HiTS-CL into five representative TKGR backbones and evaluate it on four benchmark datasets. HiTS-CL consistently improves extrapolation accuracy, reduces long-horizon degradation, and outperforms strong continual-learning baselines, including a recent method for temporal knowledge graphs. Source code and data are available at https://github.com/liuyansong98/HiTS-CL.
Figures & tables
Figure 1 . (a) Long-horizon performance decay of a fixed-prefix backbone versus HiTS-CL on the test stream (MRR averaged over 100 snapshots). Dashed/solid curves denote the backbone and its HiTS-CL variant. (b) Fixed-prefix training versus continual updating evaluated on the immediately following test snapshot after the training period (snapshot 1001; train: 1–1000).
Figure 2 . Overview of HiTS-CL: continual fine-tuning captures current dynamics, multi-teacher distillation preserves stable knowledge while incorporating time-varying patterns, and selective-memory history enhancement provides recurring historical evidence.
Dataset
Entities
Relations
Snapshots
Train
Valid
Test
Test Snapshots
New Entities
Granularity
ICEWS14
7,218
230
365
27,391
4,514
58,825
231
2,149
24 hours
ICEWS18
23,033
256
304
140,714
23,320
304,524
203
5,082
24 hours
ICEWS05-15
10,488
251
4,017
138,576
22,959
299,794
2,725
3,289
24 hours
GDELT
7,691
240
2,751
683,646
114,076
1,480,683
1,688
1,304
15 mins
Table 1 . Statistics of the datasets. All facts are chronologically split into training, validation, and test sets.
Model
ICEWS14
ICEWS18
ICEWS05-15
GDELT
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
RE-GCN (2021)
33.81
25.66
37.48
49.47
28.41
19.33
32.00
46.23
44.29
34.76
49.07
62.65
18.41
11.56
19.45
31.76
RE-GCN+FT
40.00
30.91
44.61
57.13
31.16
21.48
35.22
50.12
49.75
39.51
55.35
69.18
21.79
13.94
23.56
37.13
RE-GCN+Replay
41.29
32.16
45.86
58.61
32.48
22.53
36.75
51.93
51.64
41.21
57.42
71.31
21.75
13.86
23.46
37.21
RE-GCN+Reg
40.10
31.10
44.55
57.02
31.67
21.75
35.77
51.23
52.38
41.99
58.09
72.12
21.10
13.42
22.72
36.09
RE-GCN+HiTS-CL
43.61
34.61
48.48
60.13
34.34
24.44
38.78
53.52
52.53
42.45
58.18
71.56
25.22
16.51
27.83
42.28
Table 2 . Time-aware filtered MRR and Hits@K on ICEWS14/18, ICEWS05-15, and GDELT. We report each backbone under fixed-prefix training and continual learning variants (+FT, +Replay, +Reg, +HiTS-CL).
Figure 3 . MRR degradation trends over time of ICEWS05-15.
Figure 4 . Comparison of DGAR and LogCL+HiTS-CL.
Figure 5 . Fixed-prefix training versus continual updating under the same evaluation snapshot (MRR).
Model
ICEWS14
ICEWS18
ICEWS05-15
GDELT
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
RE-GCN w/o St.1
42.11
33.46
46.61
58.04
33.69
24.01
38.05
52.45
51.23
41.55
56.71
69.28
23.91
15.44
26.36
40.51
RE-GCN w/o St.2
42.47
33.60
47.33
58.74
33.54
23.83
37.91
52.32
51.59
41.51
57.24
70.58
23.09
16.43
27.69
42.03
RE-GCN w/o His.
41.52
32.36
46.19
58.60
32.20
22.36
36.31
51.54
51.19
41.03
56.69
70.46
22.07
14.14
23.85
37.59
RE-GCN+HiTS-CL
43.61
34.61
48.48
60.13
34.34
24.44
38.78
53.52
52.53
42.45
58.18
71.56
25.22
16.51
27.83
42.28
CEN w/o St.1
39.58
31.10
44.33
54.94
31.95
22.20
36.33
50.93
46.80
37.38
52.24
64.22
21.07
13.05
22.93
36.86
Table 3 . Ablation study on ICEWS14, ICEWS18, ICEWS05-15 and GDELT. All metrics are time-aware filtered.
Backbone
Validation
ICEWS14
ICEWS05-15
GDELT
ICEWS18
DiMNet
Gt−1
41.50
53.08
24.61
32.82
Gt (15%)
41.03
52.49
24.36
32.69
RE-GCN
Gt−1
43.61
52.53
25.22
34.34
Gt (15%)
43.29
51.48
25.14
33.43
Table 4 . Effect of validation strategy on two backbone models (MRR). Using Gt−1 for validation consistently performs slightly better than holding out 15% Gt .
ICEWS14
ICEWS05-15
ICEWS18
GDELT
Runtime Per snapshot (s)
0.064
0.032
0.408
0.297
Memory usage (MB)
134.18
339.35
535.66
965.93
Table 5 . Inference-time computational cost of the History Enhancement (His) module. We report the average runtime per snapshot and the memory usage of the historical cache on four datasets.
Figure 6 . Runtime comparison (in seconds) between full training and continual learning.
Model
Pretrain
Fine-Tune
Distillation
Overall
DiMNet
620
279
492
1,391
LogCL
1,618
736
1,153
3,507
RE-GCN
526
183
527
1,236
Table 6 . Training time breakdown of HiTS-CL (seconds) on ICEWS14, including pretraining, continual fine-tuning, and distillation stages.
Query
(Media_(China), Make_statement, ?, 299)
HiTS-CL w/o His. (Top-5)
Japan, South_Korea, Kim_Jong-Un, China , North_Korea
HiTS-CL (Top-5)
China , Japan, South_Korea, Kim_Jong-Un, North_Korea
Historical Entities
China (freq=11, latest=261) , Japan (freq=1, latest=222)
Table 7 . Case studies illustrating the effect of the History Enhancement module. Entities in bold denote the ground-truth entity.
Figure 7 . Sensitivity analysis of α and β hyper-parameters.
Temporal knowledge graph (TKG) extrapolation is an important task that aims to predict future facts through historical interaction information within KG snapshots. A key challenge for most existing TKG extrapolation models is handling entities with sparse historical interaction. The ontological knowledge is beneficial for alleviating this sparsity issue by enabling these entities to inherit behavioral patterns from other entities with the same concept, which is ignored by previous studies. In this paper, we propose a novel encoder-decoder framework OntoTKGE that leverages the ontological knowledge from the ontology-view KG (i.e., a KG modeling hierarchical relations among abstract concepts as well as the connections between concepts and entities) to guide the TKG extrapolation model's learning process through the effective integration of the ontological and temporal knowledge, thereby enhancing entity embeddings. OntoTKGE is flexible enough to adapt to many TKG extrapolation models. Extensive experiments on five data sets demonstrate that OntoTKGE not only significantly improves the performance of many TKG extrapolation models but also surpasses many state-of-the-art(SOTA) baseline methods.
Temporal Knowledge Graphs (TKGs) record how facts evolve over time, but forecasting future events on a TKG remains difficult for three reasons: (i) long-range temporal dependencies are hard to encode; (ii) events on different chains mutually excite or inhibit one another in ways that snapshot-level models cannot express; and (iii) inter-arrival times are heavy-tailed and statistically sparse, so deterministic time predictors are unreliable. We address these three issues with a single framework, the \textbf{Group Attention Neural Hawkes Process (GAttNHP)}, built around three matched components. First, a self-attention encoder casts each subject--relation chain as a continuous-time point process and captures the lingering excitation of distant history. Second, a semantic soft-grouping module turns globally learnable Hawkes priors into an analytical cross-attention mask, so chains share excitation patterns through their latent group memberships rather than through exhaustive pairwise computation. Third, a Non-Crossing Quantile (NCQ) regression head replaces mean-based time prediction, providing calibrated, monotonically ordered quantile estimates that remain stable under heavy-tailed inter-arrival distributions. On six benchmark TKG datasets, GAttNHP improves over state-of-the-art baselines on both entity prediction and time prediction, and ablations confirm that its largest gains arise on the long-tail event chains where existing models fail most severely.
Xiangni Tian, Kaixian Yu, Runpeng Dai +2
Yunnan Key Laboratory of Statistical Modeling and Data Analysis, Yunnan University, Kunming, China · Insilicom LLC, Peabody, MA, USA · University of North Carolina at Chapel Hill, North Carolina, USA
Temporal knowledge graphs (TKGs) represent time-stamped relational facts and support a wide range of reasoning tasks over evolving events. However, existing methods produce entity representations that are static at the entity level, in that each representation is a function of learned parameters only and retains no trace of the interactions in which the entity has participated. In this paper, we depart from this static view and propose that each entity be modeled as an adaptive process whose representation is refined every time the entity participates in a fact. To this end, we propose AdaTKG, which maintains a per-entity memory that is updated with every observed interaction, with the memory accumulating online and predictions improving as more interactions arrive. Specifically, we instantiate the memory update as a learnable exponential moving average governed by a single shared scalar instead of using learnable parameters for each entity, enabling AdaTKG to handle entities unseen during training. Extensive experiments confirm consistent gains over TKG baselines, demonstrating the effectiveness of adaptive memory. Code is available at: https://github.com/seunghan96/AdaTKG