Class-Incremental Learning (CIL) aims to continually learn new classes while preserving prior knowledge. Parameter-efficient fine-tuning with pre-trained models enables CIL with minimal parameter updates, but existing approaches still suffer from catastrophic forgetting caused by cumulative interference and suboptimal module-sample matching at inference. We propose Hyperbolic Prototype Routing (HyPro), a rehearsal-free framework for continual learning. HyPro allocates a dedicated LoRA-Expert module to each incremental task for isolated representation learning, then projects routing features onto a Poincare ball and performs geodesic nearest-prototype matching for reliable task-level discrimination. Extensive experiments on standard CIL and Few-Shot CIL benchmarks show that HyPro consistently improves average and final-stage accuracy over strong baselines.
Figures & tables
Fig. 1 : Illustration of HyPro. In the t -th incremental task, a new LoRA-Expert Et (with parameters At and Bt ) is trained to capture task-specific features. Domain-specific features from the router Erouter are projected into the hyperbolic space (Poincaré ball). During inference, the nearest prototype guides LoRA-Experts selection for each input sample.
Method
CIFAR100 ( T =10)
CUB200 ( T =10)
AL
Aˉ
AL
Aˉ
Full Fine-Tuning
66.26
76.94
55.29
70.30
SimpleCIL [ 7 ]
81.27
87.13
82.28
91.85
L2P [ 5 ]
84.82
89.78
71.98
81.80
CODA-Prompt [ 6 ]
86.69
91.31
75.45
84.65
InfLoRA [ 8 ]
86.43
91.80
70.07
81.71
TABLE I : Performance comparison of selected CIL methods, all built on the same pre-trained backbone ( ViT-B/16-IN21K ).
Method
CUB200 ( T =11)
CIFAR100 ( T =9)
ABase
AL
Aˉ
ABase
AL
Aˉ
L2P [ 5 ]
91.50
50.04
66.70
93.43
55.75
71.81
CODA-Prompt [ 6 ]
91.50
53.65
69.30
94.05
57.10
73.11
InfLoRA [ 8 ]
92.45
45.18
66.27
94.92
57.41
74.28
SD-LoRA [ 9 ]
91.92
56.28
70.87
94.60
73.51
78.42
CPE-CLIP [ 32 ]
80.21
63.32
69.37
88.32
79.99
83.38
TABLE II : Performance comparison of selected FSCIL methods, all built on the same pre-trained backbone ( ViT-B/16-IN21K ).
Ablated Components
ImageNet-R ( T =5)
CIFAR100 ( T =9)
AL
Aˉ
AL
Aˉ
w/o LoRA Dynamically
61.17 / 72.37
74.31 / 80.14
73.03 / 79.73
80.39 / 86.59
w/o HPR
69.53 / 73.13
77.75 / 80.42
81.13 / 84.09
86.97 / 88.67
HyPro-MLP / QV
77.00 / 78.10
82.44 / 83.05
88.96 / 88.08
91.16 / 90.54
TABLE III : Ablation studies on CIL and FSCIL tasks. The first dataset corresponds to CIL, and the second to FSCIL. For each metric, the left/right values represent performance with MLP-LoRA and QV-LoRA fine-tuning, respectively. HPR: Hyperbolic Prototype Routing.
Method
CIFAR100 ( T =10)
CUB200 ( T =10)
ImageNet-R ( T =5)
KNN
86.80
90.40
75.24
Prototype
89.60
91.10
77.04
HyPro-MLP
93.71
93.40
88.48
HyPro-QV
94.13
93.32
88.61
TABLE IV : Router average accuracy comparison of different module-sample matching strategies on CIL tasks. All methods are based on the same pre-trained backbone ( ViT-B/16-IN21K ).
Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. Parameter-efficient fine-tuning with pre-trained models reduces parameter overhead but can suffer from cumulative interference and suboptimal alignment between inference samples and specialized modules. We propose Dynamic LoRA-Experts and Prototype-Ensemble Matching (DLEPEM), a two-stage rehearsal-free framework. DLEPEM allocates a task-specific LoRA-Expert for each incremental task to reduce cross-task interference, then combines frozen pre-trained-model prototypes with task-adaptive LoRA-Expert prototypes for reliable task-level discrimination. Experiments on standard CIL and Few-Shot CIL benchmarks demonstrate strong performance under the evaluated protocols.
Hongwei Zhao, Rui Liu, Yansong Liu
School of Computer Science and Engineering, Beihang University
Class Incremental Learning (CIL) aims to learn new concepts consistently from a data stream without forgetting. Unlike typical CIL methods which need to learn a model from scratch, pre-trained model (PTM) can easily adapt to a new task with fine-tuning. However, existing PTM-based CIL methods fail to achieve a trade-off between performance and computational expenditure, i.e., they either adopt the same parameter space so that leading catastrophic forgetting, or expand a new branch for each task but adding more computational cost. To this end, we propose MetrIc Learning with Expandable Subspace (Miles) to harness the prior information within pre-trained knowledge, thereby orchestrating an efficient expansion of the parameter space through guided optimization. Specifically, it decouples the learnable modules with the pre-trained model, exploiting prior information from intermediate features of the backbone network to enable more flexible parameter expansion. Then, a central loss is adopted to guide the new category to cluster towards the corresponding prototype in the new task subspace while incorporating an auxiliary distance regularization term to maintain metric equilibrium across tasks. Extensive experiments on six benchmark datasets demonstrate that Miles achieves state-of-the-art performance in various CIL settings.
Kai Jiang, Zisong Lin, Hongyuan Zhang +2
National Key Laboratory of Radar Signal Processing, Xidian University, Xi’an 710071, China · The University of Hong Kong, Hong Kong SAR, China · Institute of Artificial Intelligence (TeleAI) of China Telecom, China
We present HydraCIL, a decoupled continual learning model based on prototype-guided multi-head classifiers, targeting sustainable deployment in embedded and resource-constrained environments. While most Class-Incremental Learning (CIL) methods rely on powerful hardware and long retraining cycles, real-world systems, such as robots or edge AI devices, must adapt quickly with limited resources. HydraCIL addresses this gap by freezing the backbone and decoupling feature extraction from learning. For each task, features are extracted once and a lightweight, task-specific classifier head is created, avoiding costly backbone retraining. At inference, HydraCIL selects the appropriate head via similarity with prototypes. Experiments on CIFAR-100, ImageNet-100, CoRe50, and Flowers102 datasets show that HydraCIL matches or outperforms state-of-the-art CIL methods while significantly reducing training time and carbon footprint, making it a practical solution for continual learning in real-world and embedded settings, where energy efficiency and rapid adaptation are critical.
Daniel Vila-Cruz, Laura Morán-Fernández, Verónica Bolón-Canedo