Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all classes seen so far without access to its task identity. For learners equipped with task-specific parameter-efficient modules, this introduces a critical task-routing challenge beyond catastrophic forgetting. We study post-hoc task routing without retraining the learner or introducing a separately trained router. Such training-free inference-time calibration remains comparatively underexplored in parameter-efficient class-incremental learning. We identify three sources of routing error (feature-level, task-level, and class-level misalignment) and propose Feature Distribution Calibration (FDC). Its three components address these misalignments: Task Subspace Filtering (TSF) suppresses feature components outside each task's principal subspace, Residual Likelihood Calibration (RLC) evaluates the typicality of its subspace residual, and Prototype Affinity Calibration (PAC) measures compatibility with the task's class prototypes. Experiments demonstrate plug-and-play applicability to eight parameter-efficient class-incremental methods using a shared encoder. With one component configuration selected per method across all five benchmarks, FDC improves final accuracy in all 40 method-dataset pairs by 4.39 percentage points on average. Enabling all components improves 35 of the 40 pairs, with an average gain of 4.45 points. When applied to a simple baseline, FDC achieves strong overall performance.
Figures & tables
Figure 1: Post-hoc task routing with FDC . Top: task-local heads lack foreign-task negatives; H∗ illustrates an ideal global classifier (Section 3 ). Bottom: TSF, RLC, and PAC calibrate routing through feature support, task membership, and class affinity, respectively, without additional training.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Classes
Tasks
Classes/task
Train
Test
CIFAR-100
100
20
5
50,000
10,000
ImageNet-A
200
10
20
5,981
1,519
ImageNet-R
200
10
20
24,000
6,000
CUB-200
200
20
10
9,430
2,358
OmniBenchmark
300
10
30
89,697
5,985
Appendix
Table 4: Dataset partitions and task sequences.
Setting
Value
Main-training epochs per task / batch size
20 / 64
Optimizer / weight decay
Adam / 0
LoRA learning rate / head learning rate
5×10−4 / 5×10−3
LoRA rank / adapted projections
10 / key and value
Training seeds / class-order seed
1,2,3 / 1993
PCA retained energy η
0.75
Appendix
Table 5: Common training and FDC hyperparameters.
Time (min)
Peak GPU memory (GiB)
FDC storage
Dataset
Baseline
+ FDC
Baseline
+ FDC
(MiB)
CIFAR-100
132.2
+5.1
6.51
+0.00
11.86
ImageNet-A
18.7
+1.1
6.51
+0.00
10.05
ImageNet-R
64.8
+5.1
6.51
+0.00
11.30
CUB-200
29.3
+3.0
6.51
+0.00
7.68
OmniB.
239.3
+9.2
6.51
+0.00
12.03
Appendix
Table 6: Training-stage time, GPU memory, and FDC storage.
Total latency (ms/batch)
Peak GPU memory (MiB)
Dataset
Baseline
Single
Dual
Baseline
Δ Single
Δ Dual
CIFAR-100
170.5
173.3 (+1.7%)
323.1 (+89.5%)
980.9
+4.7
+11.9
ImageNet-A
171.0
172.5 (+0.8%)
321.5 (+88.0%)
981.4
+5.0
+10.1
ImageNet-R
170.5
171.7 (+0.7%)
320.1 (+87.7%)
981.4
+5.9
+11.3
CUB-200
171.3
174.1 (+1.6%)
324.2 (+89.3%)
981.4
+3.7
+7.8
OmniB.
169.9
171.3 (+0.8%)
319.2 (+87.9%)
982.0
+5.7
+12.1
Appendix
Table 7: Inference latency and GPU memory at batch size 64.
Class-incremental learning (CIL) requires a model to incrementally learn tasks that contain new classes without accessing earlier training data while preserving the ability to recognize all seen classes. Recently, pretrained-model-based approaches have become prevalent by adapting a frozen backbone with additional lightweight trainable modules. Existing methods, however, exhibit limitations: task-specific adapters learn explicit per-task representations but are parameter- and computation-inefficient, while LoRA-based merging methods combine per-task LoRA parameters into a single model whose static aggregated weights cause representation interference during inference. To address these problems, we present \textbf{FACET}: task-conditioned \textbf{F}e\textbf{A}ture transformation with \textbf{C}ondition\textbf{E}d feature consis\textbf{T}ency, achieving excellent parameter efficiency while producing highly discriminative features during inference. When continually trained on a task sequence, FACET learns a single shared adapter that employs a dynamic task-conditioned feature transformation, shaping the overall feature distribution of the adapter into a mixture of overlap-reduced task-specific components. On the other hand, we propose an efficient replay-free task-conditioned feature consistency loss, aiming to mitigate catastrophic forgetting of the learned mixture distribution in the adapter's feature space. Even when maintaining only a single adapter, FACET demonstrates robust scalability. On both very long task sequences (e.g., 200 tasks) and standard short task sequences (e.g., 20 tasks), our method achieves superior performance while using significantly fewer trainable parameters and GFLOPs. The code will be made open source upon acceptance.
Yunxiang Fu, Meng Lou, Yizhou Yu
School of Computing and Data Science The University of Hong Kong
Few-Shot Class-Incremental Learning (FSCIL) addresses the challenge of learning new classes from very limited samples while retaining knowledge of previously learned ones. Although parameter-efficient fine-tuning methods with pre-trained models show promise for class-incremental learning, strict gradient-based constraints can be unreliable under severe data scarcity, while multi-expert approaches can impose substantial inference-time costs. We propose TALON (Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer), an inference-efficient FSCIL framework. TALON dynamically allocates an independent LoRA-Teacher to each incremental task for task-specific representation learning, then distills multiple frozen teachers into a unified LoRA-Student through Ensemble Knowledge Transfer, eliminating runtime module selection or generation. A semantic-guided distillation strategy weights teacher contributions by feature-space similarity to mitigate catastrophic forgetting and overfitting. Across three class-order runs, TALON achieves comparable or better mean average accuracy across four FSCIL benchmarks, obtaining 86.68 +/- 1.22% on CUB200, 90.39 +/- 0.27% on CIFAR100, 78.38 +/- 0.94% on ImageNet-R, and 96.34 +/- 0.33% on miniImageNet. TALON uses up to 33x fewer deployment parameters and reduces average inference time per task to 26.7 s, a 41.70% reduction relative to ASP.
Hongwei Zhao, Rui Liu, Yansong Liu +2
School of Computer Science and Engineering, Beihang University · School of Computer Science, Beijing University of Posts and Telecommunications
Class-Incremental Learning (CIL) aims to continually learn new classes while preserving prior knowledge. Parameter-efficient fine-tuning with pre-trained models enables CIL with minimal parameter updates, but existing approaches still suffer from catastrophic forgetting caused by cumulative interference and suboptimal module-sample matching at inference. We propose Hyperbolic Prototype Routing (HyPro), a rehearsal-free framework for continual learning. HyPro allocates a dedicated LoRA-Expert module to each incremental task for isolated representation learning, then projects routing features onto a Poincare ball and performs geodesic nearest-prototype matching for reliable task-level discrimination. Extensive experiments on standard CIL and Few-Shot CIL benchmarks show that HyPro consistently improves average and final-stage accuracy over strong baselines.
HongWei Zhao, Rui Liu, Yong Chen
Beihang University · Beijing University of Posts and Telecommunications