Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all classes seen so far without access to its task identity. For learners equipped with task-specific parameter-efficient modules, this introduces a critical task-routing challenge beyond catastrophic forgetting. We study post-hoc task routing without retraining the learner or introducing a separately trained router. Such training-free inference-time calibration remains comparatively underexplored in parameter-efficient class-incremental learning. We identify three sources of routing error (feature-level, task-level, and class-level misalignment) and propose Feature Distribution Calibration (FDC). Its three components address these misalignments: Task Subspace Filtering (TSF) suppresses feature components outside each task's principal subspace, Residual Likelihood Calibration (RLC) evaluates the typicality of its subspace residual, and Prototype Affinity Calibration (PAC) measures compatibility with the task's class prototypes. Experiments demonstrate plug-and-play applicability to eight parameter-efficient class-incremental methods using a shared encoder. With one component configuration selected per method across all five benchmarks, FDC improves final accuracy in all 40 method-dataset pairs by 4.39 percentage points on average. Enabling all components improves 35 of the 40 pairs, with an average gain of 4.45 points. When applied to a simple baseline, FDC achieves strong overall performance.
Figures & tables
Figure 1: Post-hoc task routing with FDC . Top: task-local heads lack foreign-task negatives; H∗ illustrates an ideal global classifier (Section 3 ). Bottom: TSF, RLC, and PAC calibrate routing through feature support, task membership, and class affinity, respectively, without additional training.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Classes
Tasks
Classes/task
Train
Test
CIFAR-100
100
20
5
50,000
10,000
ImageNet-A
200
10
20
5,981
1,519
ImageNet-R
200
10
20
24,000
6,000
CUB-200
200
20
10
9,430
2,358
OmniBenchmark
300
10
30
89,697
5,985
Appendix
Table 4: Dataset partitions and task sequences.
Setting
Value
Main-training epochs per task / batch size
20 / 64
Optimizer / weight decay
Adam / 0
LoRA learning rate / head learning rate
5×10−4 / 5×10−3
LoRA rank / adapted projections
10 / key and value
Training seeds / class-order seed
1,2,3 / 1993
PCA retained energy η
0.75
Appendix
Table 5: Common training and FDC hyperparameters.
Time (min)
Peak GPU memory (GiB)
FDC storage
Dataset
Baseline
+ FDC
Baseline
+ FDC
(MiB)
CIFAR-100
132.2
+5.1
6.51
+0.00
11.86
ImageNet-A
18.7
+1.1
6.51
+0.00
10.05
ImageNet-R
64.8
+5.1
6.51
+0.00
11.30
CUB-200
29.3
+3.0
6.51
+0.00
7.68
OmniB.
239.3
+9.2
6.51
+0.00
12.03
Appendix
Table 6: Training-stage time, GPU memory, and FDC storage.
Total latency (ms/batch)
Peak GPU memory (MiB)
Dataset
Baseline
Single
Dual
Baseline
Δ Single
Δ Dual
CIFAR-100
170.5
173.3 (+1.7%)
323.1 (+89.5%)
980.9
+4.7
+11.9
ImageNet-A
171.0
172.5 (+0.8%)
321.5 (+88.0%)
981.4
+5.0
+10.1
ImageNet-R
170.5
171.7 (+0.7%)
320.1 (+87.7%)
981.4
+5.9
+11.3
CUB-200
171.3
174.1 (+1.6%)
324.2 (+89.3%)
981.4
+3.7
+7.8
OmniB.
169.9
171.3 (+0.8%)
319.2 (+87.9%)
982.0
+5.7
+12.1
Appendix
Table 7: Inference latency and GPU memory at batch size 64.