Temporal Heterogeneous Graph Pretraining for Relational Deep Learning
Organizations: RWTH Aachen University, Aachen, Germany · Fraunhofer FIT, Sankt Augustin, Germany
Abstract
Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work often treats time as a single signal or studies temporal representation and pretraining separately. We investigate how explicitly encoding both signals affects temporal pretraining for downstream tasks. Our framework combines Multi-scale Time Encoding, which captures record age using learnable time scales and type-specific projections, with Rotary Time Encoding, which represents signed inter-record intervals through rotary transformations during graph propagation. We pair these encodings with three self-supervised objectives: historical relation recovery, horizon-aware future relation activity prediction, and temporal subgraph contrast. All inputs respect their observation cutoffs. Pretraining proceeds in two stages: subgraph contrast first learns neighborhood representations, followed by refinement through either relation recovery or future activity prediction. We evaluate on five RelBench datasets across 11 classification and regression tasks using heterogeneous GNN and graph Transformer backbones. With both encodings, the best evaluated staged schedules improve over supervised training with the same encodings by 3.02% and 1.06% on the two backbones, respectively, and over controls without pretraining or either encoding by 3.24% and 2.37%.
Figures & tables
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Database | Domain | Tables | Rows | FK relations | Tasks |
|---|---|---|---|---|---|
| Arxiv | Scholarly publications | 6 | 2,733,846 | 6 | 2 |
| Avito | Online advertising | 8 | 24,653,915 | 11 | 3 |
| F1 | Motor racing | 9 | 97,606 | 13 | 2 |
| H&M | Fashion retail | 3 | 16,931,173 | 2 | 2 |
| Event | Social events | 5 | 44,822,992 | 7 | 2 |
| Database | Task | Type | Train | Validation | Test | Entities |
|---|---|---|---|---|---|---|
| Arxiv | author-category | Multiclass classification | 210,769 | 39,015 | 39,655 | 126,219 |
| paper-citation | Binary classification | 534,233 | 155,845 | 193,696 | 193,696 | |
| Avito | user-clicks | Binary classification | 59,454 | 21,183 | 47,996 | 66,449 |
| user-visits | Binary classification | 86,619 | 29,979 | 36,129 | 63,405 | |
| ad-ctr | Regression | 5,100 | 1,766 | 1,816 | 4,997 | |
| F1 | driver-dnf | Binary classification | 11,411 | 566 | 702 | 821 |
| GNN | GT | |||||
| Task | Metric | HeteroGNN | RelGNN | HGT | RelGT | THGFM |
| arxiv/author-category | Acc. | |||||
| arxiv/paper-citation | AUC | |||||
| avito/user-clicks | AUC | |||||
| avito/user-visits | AUC | |||||
| f1/driver-dnf | AUC | |||||
| Multiclass classification | Binary classification | Regression | |||||||||||||
| Encoding | Author | Mean task gain | Citation | Ignore | DNF | Churn | Clicks | Visits | Mean task gain | Attend. | Position | Sales | CTR | Mean task gain | Mean type gain |
| Acc. | AUC | AUC | AUC | AUC | AUC | AUC | MAE | MAE | MAE | MAE | |||||
| Panel A: HeteroGNN ( GNN backbone) | |||||||||||||||
| DayPE (reference) | |||||||||||||||
| TimeMix | |||||||||||||||
| DayPE + TimeRoPE | |||||||||||||||
| Pretraining strategy | Multiclass classification | Binary classification | Regression | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Temporal | Graph-level | Local stage | Author | Task gain | Citation | Ignore | DNF | Churn | Clicks | Visits | Mean task gain | Attend. | Position | Sales | CTR | Mean task gain | Mean type gain |
| Acc. | AUC | AUC | AUC | AUC | AUC | AUC | MAE | MAE | MAE | MAE | |||||||
| Panel A: DayPE without TimeRoPE . All columns use Ref. A0 below. | |||||||||||||||||
| w/o | No pretraining ( Ref. A0 ) | ||||||||||||||||
| Baseline pretraining strategies | |||||||||||||||||
| w/o | – | Un-SAGE | |||||||||||||||
| Pretraining strategy | Multiclass classification | Binary classification | Regression | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Temporal | Graph-level | Local stage | Author | Task gain | Citation | Ignore | DNF | Churn | Clicks | Visits | Mean task gain | Attend. | Position | Sales | CTR | Mean task gain | Mean type gain |
| Acc. | AUC | AUC | AUC | AUC | AUC | AUC | MAE | MAE | MAE | MAE | |||||||
| Panel A: DayPE without TimeRoPE . All columns use Ref. A0 below. | |||||||||||||||||
| w/o | No pretraining ( Ref. A0 ) | ||||||||||||||||
| Baseline pretraining strategies | |||||||||||||||||
| w/o | – | Un-SAGE | |||||||||||||||
| Attribute type | PyTorch Frame encoder | Operation |
|---|---|---|
| Categorical | EmbeddingEncoder | Learned category lookup. |
| Numerical | LinearEncoder | Standardize using column statistics, then apply a learned affine map. Missing values are imputed with the column mean. |
| Multicategorical | MultiCategoricalEmbeddingEncoder | Average learned embeddings of the categories in a cell. |
| Vector / text embedding | LinearEmbeddingEncoder | Apply a learned affine map to the input vector. |
| Timestamp | TimestampEncoder | Encode the year positionally and other calendar fields cyclically, then apply a learned linear map with bias. |
| Setting | GNN ( HeteroGNN ) | GT ( THGFM ) |
|---|---|---|
| Graph layers / hidden width | 3 / 128 | 3 / 128 |
| Table feature encoder | 4-layer ResNet, width 128 | 4-layer ResNet, width 128 |
| Table-encoder dropout | 0.2 | 0.2 |
| Text embedding / learned mapping | Fixed GloVe / | Fixed GloVe / |
| Heads per attention branch | Not applicable | 8 (16 dimensions each) |
| Graph-operator dropout | None | 0.05 |
| Setting | GNN and GT |
|---|---|
| Optimizer / learning rate / weight decay | AdamW / / |
| Learning-rate schedule | Constant; no warmup or decay |
| Single-stage pretraining epochs / steps per epoch | 10 / 10,000 |
| Two-stage pretraining epochs / steps per epoch (each stage) | 10 / 5,000 |
| Total pretraining steps (either schedule) | 100,000 |
| Pretraining batch size | 256; 512 for the H&M curricula |