Hierarchical Response Preservation for Continual Adaptation of Zero-Shot Graph-Text Models
Organizations: School of Computer Science and Engineering, Beihang University Beijing, China · School of Electronic and Information Engineering Beihang University Beijing, China · School of Software, Beihang University Beijing, China
Abstract
Pretrained graph-text models align graph representations with textual semantics, enabling recognition of unseen classes and transfer across graph domains. However, as graph data and classes continually arrive, models should learn from new supervision while retaining their zero-shot transfer capabilities and historical task knowledge. Two challenges arise: (i) new classes can overturn historical predictions despite preserved distinctions among historical classes, and (ii) overly strict response preservation can stall learning of new classes. To address these challenges, we propose Hierarchical Response Preservation (HiRP). HiRP represents this competition through a hierarchical response that keeps each historical-class probability and sums new-class probabilities, preserving historical distinctions and aggregate competition while allowing distinctions within the new class group to adapt. It further uses the geometry induced by this response to guide constrained updates, retaining useful adaptation directions while controlling response drift. Across three class-incremental settings, HiRP achieves absolute gains of 1.84-7.95 percentage points in average accuracy over the strongest compared baseline in each setting, while mitigating zero-shot transfer degradation.
Figures & tables
| Method | 4 datasets / 4 tasks | 6 datasets / 6 tasks | 4 datasets / 10 tasks | ||||||
|---|---|---|---|---|---|---|---|---|---|
| FA | AA | AF | FA | AA | AF | FA | AA | AF | |
| GCN | |||||||||
| SGD | |||||||||
| LwF | |||||||||
| EWC | |||||||||
| MAS | |||||||||
| Method | 4 datasets / 4 tasks | 6 datasets / 6 tasks | 4 datasets / 10 tasks | ||||||
|---|---|---|---|---|---|---|---|---|---|
| FA | AA | AF | FA | AA | AF | FA | AA | AF | |
| Mass + radial | |||||||||
| Mass + tangent | |||||||||
| Hier. + radial | |||||||||
| HiRP (ours) | |||||||||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Protected quantity | Inputs and enforcement |
|---|---|---|
| LwF | Historical-task outputs | Current-task inputs; distillation loss. |
| FRCL / FROMP | Function posteriors / past predictions | Inducing inputs / memorable examples; functional regularization. |
| MiB | Historical-class outputs and aggregated background/current-class probability | Segmentation inputs; background-aware distillation. |
| DKD | Target-versus-rest mass and non-target conditional distribution | Labeled distillation inputs; separately weighted losses. |
| HiRP | Historical discrimination and aggregate group competition, represented by individual historical probabilities and current-group mass | Bounded unlabeled features; local and task-level KL budgets on the same expanded candidates, with checked updates. |
| Method | AA | Zero-shot mean |
|---|---|---|
| SGD | 50.71 | 56.82 |
| Fisher proposal without KL constraints | 56.84 | 55.20 |
| HiRP (ours) | 69.74 | 58.27 |
| Dataset | Nodes | Classes | Train | Valid. | Test | Single schedule |
|---|---|---|---|---|---|---|
| Cora | 2,708 | 7 | 1,624 | 542 | 542 | |
| CiteSeer | 3,186 | 6 | 1,911 | 637 | 638 | |
| WikiCS | 11,701 | 10 | 580 | 1,769 | 5,847 | |
| Ele-Photo | 48,362 | 12 | 29,017 | 9,672 | 9,673 | |
| Sports-Fitness | 173,055 | 13 | 34,611 | 17,305 | 121,139 | – |
| Actor (HetGB) | 4,416 | 5 | 2,116 | 1,411 | 889 | – |
| Method family | Updated representation | Historical information used in adaptation |
|---|---|---|
| HiRP / four-cell ablation | Shared readout on frozen GraphCLIP features | Bounded historical and external unlabeled features; response constraints. |
| SGD / LwF | Same readout and frozen feature cache | SGD: current CE only. LwF: previous-model responses on current inputs; no historical-node CE. |
| EWC / MAS / OGD | Shared readout | Stored importance/projection information under each saved recipe; not the HiRP response constraint. |
| LoRA / G2LoRA † | Graph adapters/readout; selected graph and text adapters | Published-mechanism adaptations with their own retained state and optimization budgets. |
| InfLoRA † | Last graph-attention K/V adapters and task-wise linear heads | Activation subspaces and frozen prior factors; no historical supervised replay. |
| SD-LoRA | Last graph-attention Q/V adapters and task-wise linear heads | Frozen prior low-rank directions with learned magnitudes; no historical supervised replay. |
| Method | 4 datasets / 4 tasks | 6 datasets / 6 tasks | 4 datasets / 10 tasks |
|---|---|---|---|
| GraphGPT † | 250.9 0.8 | 742.2 128.2 | 217.5 26.4 |
| LLaGA † | 174.5 59.7 | 523.9 67.5 | 142.2 19.6 |
| InfLoRA † | 19.4 2.0 | 31.8 3.6 | 27.2 1.7 |
| SD-LoRA | 21.2 0.3 | 37.2 5.0 | 34.2 2.6 |
| HiRP (ours) | 6.7 1.1 | 16.2 5.5 | 7.2 2.5 |
| Method | Cora | CiteSeer | WikiCS | Photo | ||||
|---|---|---|---|---|---|---|---|---|
| FA | AA | FA | AA | FA | AA | FA | AA | |
| GCN | 54.61 0.18 | 31.42 0.11 | 32.45 0.31 | 31.36 0.30 | 43.76 3.69 | 36.63 4.14 | 14.58 0.04 | 16.39 0.05 |
| Cosine | 65.87 1.39 | 62.54 0.98 | 42.79 4.27 | 38.87 4.34 | 58.07 1.38 | 51.58 1.85 | 56.34 1.86 | 60.00 3.01 |
| TEEN | 59.04 4.36 | 54.27 3.70 | 54.49 2.91 | 48.21 2.67 | 59.76 1.51 | 55.31 0.60 | 49.75 3.19 | 52.89 0.99 |
| TPP-style | 68.20 0.59 | 63.53 0.47 | 70.64 0.59 | 69.10 0.55 | 58.10 1.10 | 58.61 1.23 | 50.61 0.04 | 66.57 0.09 |
| SGD | 67.65 0.28 | 56.49 0.70 | 63.85 0.50 | 60.82 0.49 | 62.22 0.06 | 55.13 0.06 | 28.68 0.05 | 28.85 0.07 |
| Method | CiteSeer | Photo | ||
|---|---|---|---|---|
| FA | AA | FA | AA | |
| BERT | 31.24 0.36 | 30.20 0.35 | 14.29 0.01 | 16.06 0.01 |
| SimpleCIL (BERT) | 57.58 4.65 | 57.75 4.23 | 31.98 6.98 | 43.63 5.02 |
| HiRP (ours) | 73.88 1.27 | 70.71 1.79 | 71.23 0.66 | 75.45 1.70 |
| Dataset | Trans. | (%) | (%) | Forgotten / certified | |
|---|---|---|---|---|---|
| Cora | 6 | 724 | 97.63 0.79 | 88.89 2.54 | 0 / 642 |
| CiteSeer | 6 | 1,107 | 98.28 1.80 | 91.82 1.91 | 0 / 1,032 |
| WikiCS | 12 | 21,565 | 98.21 0.16 | 96.42 0.24 | 0 / 20,717 |
| Photo | 15 | 44,367 | 96.41 0.18 | 80.86 0.48 | 0 / 35,906 |
| Method | 4 datasets / 4 tasks | 6 datasets / 6 tasks | 4 datasets / 10 tasks | ||||||
|---|---|---|---|---|---|---|---|---|---|
| AF | AF | AF | |||||||
| TPP-style | 68.28 0.34 | 70.57 0.34 | 2.28 0.04 | 70.79 0.06 | 72.19 0.05 | 1.40 0.02 | 59.60 0.19 | 69.58 0.24 | 9.98 0.09 |
| SimpleCIL | 59.83 1.15 | 61.59 1.19 | 1.76 0.05 | 48.57 0.73 | 49.90 0.79 | 1.33 0.07 | 50.75 0.95 | 58.51 0.87 | 7.77 0.09 |
| Hier. + radial | 66.38 0.69 | 74.17 0.63 | 7.79 0.41 | 71.53 0.33 | 76.70 0.26 | 5.17 0.17 | 67.65 0.16 | 80.42 0.12 | 12.76 0.27 |
| HiRP (ours) | 66.48 0.37 | 74.91 0.57 | 8.43 0.32 | 73.45 0.13 | 79.66 0.25 | 6.21 0.32 | 67.26 0.09 | 82.00 0.11 | 14.74 0.15 |
| Method | 4 datasets / 4 tasks | 6 datasets / 6 tasks | 4 datasets / 10 tasks | ||||||
|---|---|---|---|---|---|---|---|---|---|
| FA | AA | AF | FA | AA | AF | FA | AA | AF | |
| SGD | 75.79 0.18 | 58.35 0.22 | 25.67 0.08 | 83.02 0.14 | 58.77 0.28 | 22.70 0.10 | 50.16 0.15 | 49.96 0.15 | 41.67 0.18 |
| LwF | 76.63 0.18 | 58.43 0.61 | 25.24 0.69 | 84.14 0.07 | 61.55 0.26 | 19.06 0.18 | 56.31 0.33 | 49.64 0.30 | 41.23 0.27 |
| Memory-LwF | 77.13 0.17 | 59.37 0.63 | 24.14 0.65 | 84.99 0.08 | 63.63 0.48 | 16.75 0.36 | 54.47 0.11 | 47.39 0.39 | 44.13 0.38 |
| Memory-KD | 78.18 0.13 | 68.14 0.22 | 11.19 0.27 | 85.46 0.12 | 70.58 0.17 | 7.63 0.19 | 60.71 0.12 | 58.65 0.42 | 29.97 0.45 |
| HiRP (ours) | 80.09 0.19 | 71.12 0.26 | 8.43 0.32 | 87.04 0.42 | 73.22 0.23 | 6.21 0.32 | 75.43 0.05 | 69.86 0.13 | 14.74 0.15 |