HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
Authors: Yansong Liu, Rui Liu, Yuan Zuo, Hongwei Zhao, Da Fu, Fuwei Zhang, Fuzhen Zhuang, Yong Chen, +1 more
Organizations: Beihang University Beijing, China · Beijing University of Posts and Telecommunications Beijing, China · Hubei Engineering University Xiaogan, Hubei, China
Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over long horizons. We address this mismatch by formulating extrapolative TKGR as continual learning over streaming snapshots. Under this view, effective extrapolation must jointly handle current dynamics, stable knowledge, and recurring historical evidence. Based on these requirements, we propose History-enhanced Two-Step Continual Learning (HiTS-CL), a backbone-agnostic continual learning framework for extrapolative TKGR. HiTS-CL tracks current dynamics via continual fine-tuning, preserves stable knowledge via multi-teacher adaptive distillation, and retains recurring historical evidence via a selective memory of recent and frequent facts. We integrate HiTS-CL into five representative TKGR backbones and evaluate it on four benchmark datasets. HiTS-CL consistently improves extrapolation accuracy, reduces long-horizon degradation, and outperforms strong continual-learning baselines, including a recent method for temporal knowledge graphs. Source code and data are available at https://github.com/liuyansong98/HiTS-CL.
Figures & tables
Figure 1 . (a) Long-horizon performance decay of a fixed-prefix backbone versus HiTS-CL on the test stream (MRR averaged over 100 snapshots). Dashed/solid curves denote the backbone and its HiTS-CL variant. (b) Fixed-prefix training versus continual updating evaluated on the immediately following test snapshot after the training period (snapshot 1001; train: 1–1000).
Figure 2 . Overview of HiTS-CL: continual fine-tuning captures current dynamics, multi-teacher distillation preserves stable knowledge while incorporating time-varying patterns, and selective-memory history enhancement provides recurring historical evidence.
Dataset
Entities
Relations
Snapshots
Train
Valid
Test
Test Snapshots
New Entities
Granularity
ICEWS14
7,218
230
365
27,391
4,514
58,825
231
2,149
24 hours
ICEWS18
23,033
256
304
140,714
23,320
304,524
203
5,082
24 hours
ICEWS05-15
10,488
251
4,017
138,576
22,959
299,794
2,725
3,289
24 hours
GDELT
7,691
240
2,751
683,646
114,076
1,480,683
1,688
1,304
15 mins
Table 1 . Statistics of the datasets. All facts are chronologically split into training, validation, and test sets.
Model
ICEWS14
ICEWS18
ICEWS05-15
GDELT
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
RE-GCN (2021)
33.81
25.66
37.48
49.47
28.41
19.33
32.00
46.23
44.29
34.76
49.07
62.65
18.41
11.56
19.45
31.76
RE-GCN+FT
40.00
30.91
44.61
57.13
31.16
21.48
35.22
50.12
49.75
39.51
55.35
69.18
21.79
13.94
23.56
37.13
RE-GCN+Replay
41.29
32.16
45.86
58.61
32.48
22.53
36.75
51.93
51.64
41.21
57.42
71.31
21.75
13.86
23.46
37.21
RE-GCN+Reg
40.10
31.10
44.55
57.02
31.67
21.75
35.77
51.23
52.38
41.99
58.09
72.12
21.10
13.42
22.72
36.09
RE-GCN+HiTS-CL
43.61
34.61
48.48
60.13
34.34
24.44
38.78
53.52
52.53
42.45
58.18
71.56
25.22
16.51
27.83
42.28
Table 2 . Time-aware filtered MRR and Hits@K on ICEWS14/18, ICEWS05-15, and GDELT. We report each backbone under fixed-prefix training and continual learning variants (+FT, +Replay, +Reg, +HiTS-CL).
Figure 3 . MRR degradation trends over time of ICEWS05-15.
Figure 4 . Comparison of DGAR and LogCL+HiTS-CL.
Figure 5 . Fixed-prefix training versus continual updating under the same evaluation snapshot (MRR).
Model
ICEWS14
ICEWS18
ICEWS05-15
GDELT
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
MRR
H@1
H@3
H@10
RE-GCN w/o St.1
42.11
33.46
46.61
58.04
33.69
24.01
38.05
52.45
51.23
41.55
56.71
69.28
23.91
15.44
26.36
40.51
RE-GCN w/o St.2
42.47
33.60
47.33
58.74
33.54
23.83
37.91
52.32
51.59
41.51
57.24
70.58
23.09
16.43
27.69
42.03
RE-GCN w/o His.
41.52
32.36
46.19
58.60
32.20
22.36
36.31
51.54
51.19
41.03
56.69
70.46
22.07
14.14
23.85
37.59
RE-GCN+HiTS-CL
43.61
34.61
48.48
60.13
34.34
24.44
38.78
53.52
52.53
42.45
58.18
71.56
25.22
16.51
27.83
42.28
CEN w/o St.1
39.58
31.10
44.33
54.94
31.95
22.20
36.33
50.93
46.80
37.38
52.24
64.22
21.07
13.05
22.93
36.86
Table 3 . Ablation study on ICEWS14, ICEWS18, ICEWS05-15 and GDELT. All metrics are time-aware filtered.
Backbone
Validation
ICEWS14
ICEWS05-15
GDELT
ICEWS18
DiMNet
Gt−1
41.50
53.08
24.61
32.82
Gt (15%)
41.03
52.49
24.36
32.69
RE-GCN
Gt−1
43.61
52.53
25.22
34.34
Gt (15%)
43.29
51.48
25.14
33.43
Table 4 . Effect of validation strategy on two backbone models (MRR). Using Gt−1 for validation consistently performs slightly better than holding out 15% Gt .
ICEWS14
ICEWS05-15
ICEWS18
GDELT
Runtime Per snapshot (s)
0.064
0.032
0.408
0.297
Memory usage (MB)
134.18
339.35
535.66
965.93
Table 5 . Inference-time computational cost of the History Enhancement (His) module. We report the average runtime per snapshot and the memory usage of the historical cache on four datasets.
Figure 6 . Runtime comparison (in seconds) between full training and continual learning.
Model
Pretrain
Fine-Tune
Distillation
Overall
DiMNet
620
279
492
1,391
LogCL
1,618
736
1,153
3,507
RE-GCN
526
183
527
1,236
Table 6 . Training time breakdown of HiTS-CL (seconds) on ICEWS14, including pretraining, continual fine-tuning, and distillation stages.
Query
(Media_(China), Make_statement, ?, 299)
HiTS-CL w/o His. (Top-5)
Japan, South_Korea, Kim_Jong-Un, China , North_Korea
HiTS-CL (Top-5)
China , Japan, South_Korea, Kim_Jong-Un, North_Korea
Historical Entities
China (freq=11, latest=261) , Japan (freq=1, latest=222)
Table 7 . Case studies illustrating the effect of the History Enhancement module. Entities in bold denote the ground-truth entity.
Figure 7 . Sensitivity analysis of α and β hyper-parameters.
Yunnan Key Laboratory of Statistical Modeling and Data Analysis, Yunnan University, Kunming, China · Insilicom LLC, Peabody, MA, USA · University of North Carolina at Chapel Hill, North Carolina, USA