Organizations: School of Computer Science and Engineering, Beihang University · School of Computer Science, Beijing University of Posts and Telecommunications
Few-Shot Class-Incremental Learning (FSCIL) addresses the challenge of learning new classes from very limited samples while retaining knowledge of previously learned ones. Although parameter-efficient fine-tuning methods with pre-trained models show promise for class-incremental learning, strict gradient-based constraints can be unreliable under severe data scarcity, while multi-expert approaches can impose substantial inference-time costs. We propose TALON (Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer), an inference-efficient FSCIL framework. TALON dynamically allocates an independent LoRA-Teacher to each incremental task for task-specific representation learning, then distills multiple frozen teachers into a unified LoRA-Student through Ensemble Knowledge Transfer, eliminating runtime module selection or generation. A semantic-guided distillation strategy weights teacher contributions by feature-space similarity to mitigate catastrophic forgetting and overfitting. Across three class-order runs, TALON achieves comparable or better mean average accuracy across four FSCIL benchmarks, obtaining 86.68 +/- 1.22% on CUB200, 90.39 +/- 0.27% on CIFAR100, 78.38 +/- 0.94% on ImageNet-R, and 96.34 +/- 0.33% on miniImageNet. TALON uses up to 33x fewer deployment parameters and reduces average inference time per task to 26.7 s, a 41.70% reduction relative to ASP.
Figures & tables
Figure 1: Performance and inference time comparison among multi-expert-based and LoRA-based methods. TALON consolidates multi-expert knowledge into a single model via distillation, eliminating runtime module selection/generation overhead while achieving superior accuracy (inference time averaged per task on CIFAR100).
Figure 2: Overview of the TALON framework. (a) Task-Adaptive LoRA-Teacher: For each incremental task t , a new LoRA-Teacher Tt is trained to capture task-specific features while keeping previous LoRA-Teachers frozen. (b) Ensemble Knowledge Transfer: Multiple LoRA-Teachers are distilled into a unified LoRA-Student St using semantic-guided adaptive coefficients. (c) Semantic-Guided Distillation Coefficients αk : prototypes are extracted from the k -th LoRA-Teacher for task t data using Eq. ( 14 ), and the final coefficient αk is obtained by normalizing raw task-relevance scores in the shared feature space of LoRA-Teacher Tk according to Eq. ( 13 ).
Task
CUB200
CIFAR100
ImageNet-R
mini ImageNet
Base
100
60
100
60
Incremental
10-way 5-shot
5-way 5-shot
10-way 5-shot
5-way 5-shot
# of tasks
1+10
1+8
1+10
1+8
Table 1: Setups for the four datasets
Method
CUB200 ( T=11 )
CIFAR100 ( T=9 )
ImageNet-R ( T=11 )
mini ImageNet ( T=9 )
ABase
AL
Aˉ
ABase
AL
Aˉ
ABase
AL
Aˉ
ABase
AL
Aˉ
Full Finetune
89.36 ± 0.44
11.28 ± 3.09
22.19 ± 6.46
92.08 ± 0.54
40.86 ± 7.14
66.17 ± 4.02
82.66 ± 0.76
13.23 ± 2.93
27.99 ± 7.29
95.60 ± 0.52
69.04 ± 4.67
80.88 ± 1.67
SimpleFSCIL
88.57 ± 2.01
74.67 ± 1.41
79.71 ± 0.90
78.13 ± 1.51
63.85 ± 0.55
69.75 ± 0.69
63.73 ± 1.10
50.32 ± 0.79
55.46 ± 0.64
94.74 ± 0.66
87.55 ± 1.07
90.65 ± 0.47
L2P
90.41 ± 1.89
48.33 ± 1.06
65.13 ± 1.26
92.43 ± 0.72
55.43 ± 0.27
71.25 ± 0.38
79.79 ± 0.96
42.37 ± 2.03
57.51 ± 1.67
97.01 ± 0.30
62.09 ± 0.52
77.05 ± 0.67
CODA-Prompt
91.34 ± 1.92
50.38 ± 0.88
66.75 ± 0.68
93.50 ± 0.26
56.25 ± 0.13
72.12 ± 0.14
81.70 ± 0.36
45.67 ± 1.45
60.01 ± 1.76
97.65 ± 0.14
63.46 ± 0.91
77.68 ± 0.47
LAE
90.91 ± 2.37
50.53 ± 3.99
66.67 ± 2.85
92.19 ± 1.44
55.97 ± 2.14
71.51 ± 1.79
76.89 ± 2.89
42.79 ± 5.41
56.44 ± 4.59
97.09 ± 0.40
64.34 ± 5.56
78.29 ± 3.08
Table 2: Comparison with state-of-the-art FSCIL methods. Values are mean ± sample standard deviation over three class-order runs generated using seeds 42, 1993, and 2025. All methods use the same ViT-B/16-IN1K backbone and dataset splits. The best mean is highlighted in bold , and the second-best mean is underlined .
Method
CUB200 ( T =11)
CIFAR100 ( T =9)
ImageNet-R ( T =11)
mini ImageNet ( T =9)
PD ↓
FWT ↑
PD ↓
FWT ↑
PD ↓
FWT ↑
PD ↓
FWT ↑
Full Finetune
79.49
-17.13
51.52
34.80
66.60
15.69
31.92
9.02
SimpleFSCIL
9.96
13.67
12.74
33.80
11.56
35.89
7.32
-0.47
L2P
41.52
11.67
36.97
40.80
37.37
38.89
35.08
15.52
CODA-Prompt
39.79
17.78
37.29
44.30
36.00
45.29
33.51
13.02
LAE
42.16
9.07
36.83
45.30
36.34
35.49
38.56
-30.98
Table 3: Comparison with state-of-the-art FSCIL methods under seed 1993. We report performance drop (PD, in pp) and forward transfer (FWT) using the same ViT-B/16-IN1K backbone.
Method
A0
A1
A2
A3
A4
A5
A6
A7
A8
A9
A10
Aˉ
PD ↓
Full Finetune
88.90
2.28
3.30
7.91
6.18
9.83
10.49
13.00
11.88
7.70
9.41
15.53
79.49
SimpleFSCIL
86.25
83.23
81.69
80.07
79.33
77.70
77.30
77.26
76.48
76.38
76.29
79.27
9.96
L2P
88.81
82.91
75.74
70.03
64.65
61.23
57.40
53.46
50.69
48.41
47.29
63.69
41.52
CODA-Prompt
89.58
84.80
77.89
72.09
67.65
63.69
59.98
56.22
53.05
50.92
49.79
65.97
39.79
LAE
88.56
82.28
75.09
69.50
65.02
61.01
57.45
53.61
50.54
48.50
46.40
63.45
42.16
InfLoRA
91.12
85.43
77.96
72.09
66.24
61.75
57.88
53.82
50.64
47.92
45.46
64.57
45.66
Table 4: Detailed Top-1 accuracy (%) at each FSCIL session on CUB200 under seed 1993. We report per-session accuracy At , average accuracy Aˉ , and performance drop (PD, in pp).
Figure 3: Performance curves on different datasets. All methods are based on the same pre-trained model ( ViT-B/16-IN1K ).
Figure 4: Inference time per task (seconds) for different methods. Here, “per task” denotes one FSCIL session, and each value measures the wall-clock evaluation time at task t on all classes seen so far. Experiments are conducted on an NVIDIA A800 GPU, with batch size 48 for the base task and 16 for incremental tasks.
Figure 5: The number of expanded parameters and accuracy of different methods.
(a) Accuracy and cumulative time
Dataset
T
Aˉ (%)
Teacher train (s)
EKT time (s)
Total time (s)
CUB200
T=11
84.95
650.3±2.2
100.2±0.3
787.8±1.6
CUB200
T=21
85.24
689.8±4.1
199.5±4.4
932.4±8.4
CIFAR100
T=9
90.03
3553.0±63.3
63.6±21.2
3817.2±99.3
CIFAR100
T=21
90.10
3613.0±85.4
159.9±57.5
3981.6±159.3
ImageNet-R
T=11
77.99
1613.4±25.0
118.9±24.1
1826.2±46.6
Table 5: Performance and efficiency of TALON-QV under the dataset-specific T=11/T=9 settings and the extended T=21 setting. Training times are mean ± sample standard deviation over three repetitions (seeds 42, 1993, and 2025) on the same NVIDIA A800; peak GPU memory and state sizes are deterministic. Retained training state excludes the frozen PTM and optimizer state, while deployment retains only the LoRA-Student and Student prototype classifier.
Dataset
Method
Training params (M)
Peak GPU (MiB)
Time/epoch (s)
CUB200 ( T=11 )
ASP
2.00
5213.16
15.94
SEC-prompt
4.03
4892.90
3.63
TALON-MLP
3.10
4372.37
23.96
TALON-QV
6.19
4961.68
24.36
CIFAR100 ( T=9 )
ASP
2.00
9738.25
80.00
SEC-prompt
1.99
4747.56
18.86
Table 6: Training-phase comparison with baselines at T=11 on CUB200 and ImageNet-R and at T=9 on CIFAR100 and mini ImageNet. We report training-state expanded parameters, peak allocated GPU memory, and training time per epoch.
Dataset
Tasks
ΔAˉ
Wall increase
EKT growth
Retained growth
Gain over SEC-prompt
CUB200
11→21
+0.29
+18.4%
1.99×
1.74×
+0.78
CIFAR100
9→21
+0.07
+4.3%
2.52×
2.11×
+1.29
ImageNet-R
11→21
−0.16
+14.7%
2.33×
1.74×
+1.41
mini ImageNet
9→21
+0.03
+10.8%
2.16×
2.11×
+0.57
Table 7: Changes produced by extending TALON-QV from T=11/T=9 to T=21 . EKT and retained-state columns report multiplicative growth.
Ablated Components
CUB200 ( T =11)
CIFAR100 ( T =9)
mini ImageNet ( T =9)
AL
Aˉ
AL
Aˉ
AL
Aˉ
w/o LoRA-Teacher KD
29.13 / 71.76
45.61 / 77.50
74.49 / 85.27
79.16 / 88.08
87.70 / 93.60
93.71 / 95.20
w/o Semantic Similarity
83.14 / 83.56
84.07 / 84.02
82.12 / 86.37
88.52 / 89.19
94.12 / 94.43
95.18 / 95.42
TALON-MLP / QV
85.24 / 84.31
85.55 / 84.95
87.95 / 87.66
90.26 / 90.03
95.22 / 95.49
96.22 / 96.44
Table 8: Ablation studies of different components. For each metric, the left/right values represent performance with MLP-Teacher and QV-Teacher fine-tuning, respectively
Weighting strategy
CUB200
CIFAR100
ImageNet-R
mini ImageNet
Macro Avg.
Δ
Summed cosine (TALON-QV)
84.95
90.03
77.99
96.44
87.35
–
T0-KD
84.23
89.11
76.82
96.44
86.65
−0.70
Uniform-KD
84.16
89.27
76.16
96.46
86.51
−0.84
Maximum similarity
84.81
89.09
76.89
96.47
86.82
−0.54
Mean pairwise cosine
84.53
89.08
75.80
96.43
86.46
−0.89
Centroid cosine
84.74
89.09
75.77
96.43
86.51
−0.85
Table 9: Comparison of TALON’s summed semantic-guided coefficient with alternative task-level Teacher-weighting strategies. Values are average accuracy Aˉ (%). “Macro Avg.” is the unweighted mean over the four datasets, and Δ denotes the change relative to summed cosine. Macro averages and differences are calculated before rounding.
Method
CUB200 ( T =11)
CIFAR100 ( T =9)
ImageNet-R ( T =11)
mini ImageNet ( T =9)
ABase↑
AL↑
Aˉ↑
ABase↑
AL↑
Aˉ↑
ABase↑
AL↑
Aˉ↑
ABase↑
AL↑
Aˉ↑
Routers
88.56
79.13
81.62
93.17
57.29
72.30
84.47
55.50
64.22
97.68
93.44
95.40
Ensemble-Weights
88.73
6.91
48.48
91.28
85.57
87.94
82.86
13.53
48.77
97.52
95.14
96.03
Ensemble-Logits
90.35
67.51
76.10
94.32
64.90
77.81
87.23
50.43
64.84
97.93
85.44
91.04
TALON-MLP
89.24
85.24
85.55
93.07
87.95
90.26
84.53
72.23
77.54
97.68
95.22
96.22
TALON-QV
89.07
84.31
84.95
93.10
87.66
90.03
85.06
72.55
77.99
97.85
95.49
96.44
Table 10: Comparison of different knowledge consolidation strategies on FSCIL . We report the base, last, and average accuracy. All methods are based on the same pre-trained model ( ViT-B/16-IN1K ).
PTM
Method
ABase
AL
Aˉ
Sup-21K
Full Finetune
90.97
40.83
63.89
SimpleFSCIL
83.00
71.53
76.71
L2P
91.92
55.35
71.09
CODA-Prompt
93.23
55.89
71.87
InfLoRA
94.10
56.18
72.29
SD-LoRA
93.58
78.56
84.98
Table 11: Results of different methods on CIFAR100 ( T =9) using supervised pre-training model ViT-B/16-IN21K (denoted as Sup-21K) and self-supervised model DINO
Figure 6: Visualization of the decision boundary on mini ImageNet between two incremental tasks. Dots represent old classes, and triangles stand for new classes. Decision boundaries are shown with the shadow region.
Figure 7: Sensitivity of hyperparameters (the rank r , the insertion positions) on CUB200 ( T =11).
Dataset
Variant
Ddrift
ρΔ
Rcent
Frozen AL (%)
ΔSDC
ΔRecomp
CUB200
TALON-MLP
0.1105±0.0188
0.860
24.0%
83.90±1.08
+0.042
+0.014
TALON-QV
0.0605±0.0237
0.810
18.7%
83.80±0.85
−0.042
−0.028
CIFAR100
TALON-MLP
0.0503±0.0057
0.852
25.9%
87.42±0.66
−0.003
+0.003
TALON-QV
0.0130±0.0032
0.650
5.9%
87.48±0.23
+0.030
+0.020
ImageNet-R
TALON-MLP
0.1305±0.0279
0.874
25.7%
73.59±0.80
+0.078
−0.044
TALON-QV
0.0666±0.0040
0.766
11.4%
73.66±1.21
−0.017
−0.022
Table 12: Prototype-drift analysis and comparison of historical-prototype update strategies. Ddrift is the final-session mean cumulative L2 drift; ρΔ is the pooled drift-direction agreement; and Rcent is the pooled centroid-distance reduction. ΔSDC and ΔRecomp are mean paired changes in final accuracy relative to Frozen, in percentage points. Ddrift and Frozen AL are reported as mean ± standard deviation over class-order seeds 42, 1993, and 2025.
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China · Pengcheng Laboratory, Shenzhen, China · King Abdullah University of Science and Technology, Jeddah, Saudi Arabia +1