Organizations: School of Artificial Intelligence, Nanjing University · State Key Laboratory for Novel Software Technology, Nanjing University · Hong Kong Polytechnic University
Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features. Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VIS uses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VIS employs a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VIS accumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VIS achieves state-of-the-art performance without a textual branch.
Figures & tables
Figure 1: Preliminary diagnosis of textual embeddings in CLIP-based CIL. (a) t-SNE visualization shows a clear separation between visual prototypes and textual embeddings in CLIP’s shared image-text embedding space. (b) The class-level modality gap gc=1−μc⊤tc is positively correlated with textual-head degradation Δc=acccNCM−accctext . (c) Visual-prototype initialization yields lower training loss and higher seen-class accuracy than textual initialization under the same cosine classifier.
Figure 2: Illustration of Vis . The classifier is constructed entirely in the visual space: the frozen CLIP visual encoder extracts multi-layer CLS features, and the residual fusion module is trained with the adaptation loss Ladapt . Residual fusion and the fixed kernel feature map then produce the kernel feature ϕ(x) for classification. In the base session, Vis initializes the LS-SVM sufficient statistics (G1,Q1,s1) ; in incremental sessions (t>1) , it updates (Gt,Qt,st) with new data and recomputes the classifier weights W~t⋆ in closed form.
Method
Aircraft
CIFAR100
Cars
B0 Inc10
B50 Inc10
B0 Inc10
B50 Inc10
B0 Inc10
B50 Inc10
Aˉ
AB
Aˉ
AB
Aˉ
AB
Aˉ
AB
Aˉ
AB
Aˉ
AB
SimpleCIL Zhou et al. (2025a)
59.24
48.09
53.05
48.09
84.15
76.63
80.20
76.63
92.04
86.85
88.96
86.85
ACIL ( Zhuang et al., 2022 )
64.99
56.68
58.48
56.71
89.41
83.73
86.56
83.72
92.91
87.48
89.79
87.48
DualPrompt Wang et al. (2022b)
44.30
25.83
46.07
33.57
81.63
72.44
80.12
72.57
76.26
62.94
76.88
67.55
CODA-Prompt Smith et al. (2023)
45.98
27.69
45.14
32.28
82.43
73.43
78.69
71.58
80.21
66.47
75.06
64.19
Table 1: Main benchmark results in terms of average accuracy Aˉ and final-session accuracy AB . The best results are shown in bold. All methods start from the same pre-trained CLIP backbone for fair comparison.
Figure 3: Incremental performance of different methods. We report the performance gap after the last incremental stage of Vis and the runner-up method at the end of the line. More results are in the supplementary.
Figure 4: Ablation study, trainable-parameter comparison, and parameter sensitivity.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Φ~t←Lift(Dt;Φv,B,M,R).
Appendix
Algorithm 1 Vis task-wise update.
Figure 5: Additional analysis of Vis . Left: robustness to different class-order seeds. Middle: performance under different CLIP pre-trained weights. Right: training time comparison with representative baselines.
Method
Aircraft B0 Inc10
CIFAR100 B0 Inc10
Aˉ
AB
Aˉ
AB
ACIL ( Zhuang et al., 2022 )
64.99
56.68
89.41
83.73
RanPAC ( McDonnell et al., 2023 )
69.77
60.28
89.30
83.18
BOFA ( Li et al., 2026 )
70.96
60.43
86.07
79.19
ENGINE ( Zhou et al., 2025b )
69.74
58.51
86.89
79.29
Vis (Ours)
75.236±0.106
66.520±0.277
90.598±0.073
85.428±0.167
Appendix
Table 2: Random-feature robustness under different initialization seeds of R . Results of Vis are reported as mean ± standard deviation across random-feature seeds, while the other methods are included as references under the same incremental protocols.
Backbone
Method
Aircraft
CIFAR100
Aˉ
AB
Aˉ
AB
DINOv2 ViT-B/14
SimpleCIL ( Zhou et al., 2025a )
46.00
37.23
92.36
88.10
RanPAC ( McDonnell et al., 2023 )
87.39
79.81
90.09
83.58
Vis (Ours)
87.68
80.53
95.17
91.98
SigLIP ViT-B/16
SimpleCIL ( Zhou et al., 2025a )
78.18
69.76
81.16
74.34
RanPAC ( McDonnell et al., 2023 )
76.35
66.31
86.87
79.85
Appendix
Table 3: Generalization across different pre-trained backbones under the B0 Inc10 protocol. All methods use the same backbone within each group. We report average accuracy Aˉ and final-session accuracy AB .
Method
Food
Cars
UCF
SUN
B0 Inc10
B50 Inc10
B0 Inc10
B50 Inc10
B0 Inc10
B50 Inc10
B0 Inc30
B150 Inc30
SimpleCIL ( Zhou et al., 2025a )
6.22
3.76
5.11
2.51
4.86
2.31
8.12
3.51
ACIL ( Zhuang et al., 2022 )
4.97
2.94
5.48
2.27
1.17
0.74
8.41
4.52
DualPrompt ( Wang et al., 2022b )
24.28
20.77
7.35
4.71
9.26
7.95
23.54
17.63
CODA-Prompt ( Smith et al., 2023 )
25.82
22.41
6.14
4.59
9.80
9.21
22.71
17.69
RanPAC ( McDonnell et al., 2023 )
5.39
2.56
3.96
1.71
1.32
0.98
7.09
3.34
Appendix
Table 4: Forgetting measure FB of different methods under different incremental settings. Lower FB indicates less forgetting, and the best result in each column is highlighted in bold.
Method
Aircraft B0 Inc10
CIFAR100 B0 Inc10
Aˉ
AB
Aˉ
AB
ZS-CLIP ( Radford et al., 2021 )
26.66
17.22
81.81
71.38
Vis + Text
73.18
64.69
90.56
85.43
Vis
75.21
66.46
90.64
85.26
Appendix
Table 5: Controlled evaluation of textual information. ZS-CLIP serves as a textual-classifier reference, while Vis + Text and Vis form a controlled pair differing only in the use of textual class prototypes.
Method
Aˉ
Total Time (min) ↓
Train Time (min) ↓
Eval Time (min) ↓
Peak GPU (GB) ↓
Infer. (ms/img) ↓
DualPrompt ( Wang et al., 2022b )
63.43
48.90
45.68
3.22
6.95
8.58
CODA-Prompt ( Smith et al., 2023 )
66.81
53.23
49.64
3.59
15.53
9.58
SimpleCIL ( Zhou et al., 2025a )
92.11
1.50
0.37
1.13
0.92
3.02
RAPF ( Huang et al., 2024 )
82.12
19.96
18.62
1.34
6.43
3.59
CLG-CBM ( Yu et al., 2025 )
93.15
32.90
31.77
1.13
1.05
3.01
PROOF ( Zhou et al., 2025c )
90.44
33.26
30.85
2.41
1.16
6.43
Appendix
Table 6: Overhead comparison on Cars B0 Inc10 . All methods start from the same pre-trained CLIP backbone and are evaluated under the same hardware environment, using a single RTX 3090 GPU per method. “Total Time” denotes the end-to-end wall-clock time from the start of continual training to the completion of the final evaluation. “Peak GPU” denotes the peak training GPU memory of the current process, and “Infer.” reports the average evaluation latency per test image aggregated over all sessions. Lower is better for all overhead metrics, while higher is better for Aˉ .
Figure 6: Incremental performance of different methods on half-base setting. We report the performance gap after the last incremental stage of Vis and the runner-up method at the end of the line. All methods utilize the same CLIP pre-trained weight.
Figure 7: Incremental performance of different methods on B0 setting. We report the performance gap after the last incremental stage of Vis and the runner-up method at the end of the line. All methods utilize the same CLIP pre-trained weight.
Figure 8: Classifier-choice ablation under fixed kernel-induced visual features. All methods use the same enhanced representation and kernel feature map, and differ only in the final classifier. NCM uses class means, Ridge uses one-hot least-squares targets, and LS-SVM uses one-vs-all ±1 targets. Ridge and LS-SVM both outperform NCM, while LS-SVM achieves comparable or better performance than Ridge in most settings, supporting our SVM-style positive-versus-negative classifier formulation.
Figure 9: Cross-class gradient cosine distributions under textual initialization and image-prototype initialization on representative datasets. We compute pairwise cosine similarities cos(∇c,∇c′) between class-wise gradient directions. Compared with textual initialization, image-prototype initialization generally produces distributions that are more concentrated around zero, indicating more balanced class-wise optimization directions.
Figure 10: Training-loss curves under textual initialization and image-prototype initialization on the same six representative datasets as Figure 12 . Textual initialization often leads to larger loss values and stronger loss fluctuations, while image-prototype initialization produces lower or smoother loss curves. This suggests that visual prototypes provide a more stable initialization for classifier optimization.
Figure 11: Additional class-wise modality-gap analysis on nine datasets. Each point denotes one class. The x-axis is the class-level modality gap gc between the image prototype and textual embedding, and the y-axis is the class-wise accuracy improvement Δc of image-prototype initialization over textual initialization. The dashed red line shows the linear fit, and the Pearson correlation coefficient is reported in each subplot. Most datasets show a positive correlation, suggesting that classes with larger image-text gaps tend to benefit more from image-prototype initialization.
Figure 12: Seen-class accuracy curves under textual initialization and image-prototype initialization on six representative datasets. The x-axis denotes global training epochs, and the y-axis denotes the accuracy over all seen classes. Image-prototype initialization generally yields higher and more stable seen-class accuracy, indicating a more favorable optimization trajectory.
Figure 13: Additional t-SNE visualizations of image samples, image prototypes, and textual embeddings on nine datasets. Blue dots denote image samples, red triangles denote image prototypes, and orange stars denote textual embeddings. Textual embeddings often lie in regions separated from the corresponding image distributions, while image prototypes remain close to image samples. This illustrates the image-text modality gap in the shared CLIP embedding space.