Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.
Figures & tables
Figure 1: Motivating observations. (a) New-task cross-modal gradients retain substantial projection energy in historically occupied directions ( Kh=128 ) across the visual tower. (b) Tuning occupied directions causes greater old-class relation-structure drift than tuning low-interference directions. SR denotes rank-structure similarity to the old-task relation matrix, whose entries are cosine similarities between old-class visual prototypes and corresponding text embeddings.
Figure 2: Overview of DuLBE . (a) Dual-mode allocation uses the prospective task gradient to form a compact shared mode within historically occupied directions and a low-interference residual mode outside them. (b) Routed cross-modal training directs gradients from the current-task objective to both learners, while routing gradients from the semantic-structure preservation objective only to the shared learner. (c) Classifier: it builds visual-text geodesic bridges, estimates bridge-depth reliability after each task, and ensembles reliable bridge prototypes for inference.
Method
Setting-A (CLIP)
Setting-B (OpenCLIP)
CIFAR
CUB
I.N.-R
I.N.100
CIFAR
SUN
Cars
Food
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
ContinualCLIP
–
66.7
–
51.2
–
72.0
–
75.4
–
71.4
–
72.1
–
76.4
–
81.9
DualPrompt (ECCV’22)
81.5
72.5
–
–
82.0
75.8
80.7
67.4
81.6
72.4
82.5
74.4
76.3
62.9
84.9
77.3
CODA (CVPR’23)
77.0
62.3
66.6
50.9
78.0
67.5
64.1
34.8
82.4
73.4
83.3
75.7
80.2
66.5
86.2
78.8
PROOF (TPAMI’25)
86.2
76.3
–
–
82.8
77.1
84.7
72.5
86.8
79.1
83.9
77.3
90.7
86.5
90.0
84.7
Table 1: Main comparison under two standard CLIP-based CIL settings. Setting-A ( Huang et al. 2024 ) uses OpenAI CLIP ViT-B/16, and Setting-B ( Zhou et al. 2025a ) uses OpenCLIP ViT-B/16 pretrained on LAION-400M. All methods are evaluated using the standard 10-task split without external training data . The best results are highlighted in bold , and the second-best results are underlined .
Mode bases
∇Lc-ce
∇Lbi-kl
CIFAR (A)
CUB (A)
BS
BR
BS
BR
Aˉ↑
Al↑
Aˉ↑
Al↑
LoRA ∗ ( r=32 )
–
–
–
–
(86.1)
(78.5)
(81.0)
(73.8)
Shared PS
✓
–
✓
×
+2.4
+4.1
+3.2
+4.0
Residual PR
–
✓
×
×
+2.5
+4.4
+4.7
+5.5
Dual [PS,PR]
✓
✓
×
×
+1.3
+2.4
+1.1
+1.5
Dual [PS,PR]
✓
✓
✓
✓
+2.6
+4.4
+4.2
+5.3
Table 2: Ablation of dual-mode collaborative learning. All values are improvements over the LoRA ∗ baseline, which replaces the low-rank learner with standard LoRA while keeping all other components unchanged.
Method
Text Clf.
B.E. Clf.
CUB
Food
α
Rel.
CUB
Food
Ours
67.9
83.7
0.25
×
77.9
85.0
Ours
–
–
0.50
×
77.7
83.9
Ours
–
–
0.75
×
72.9
84.2
Ours
–
–
{αk}
×
79.1
84.6
Ours
–
–
{αk}
✓
80.6 (+12.7)
86.7 (+3.0)
Table 3: Ablation of the bridge-prototype ensemble on CUB (A) and Food (B). All entries report the last accuracy Al . B.E. denotes bridge-prototype ensemble and Rel. denotes reliability weighting over bridge depths.
Setting
Dataset
Classes
Train
Test
Stream
Training template
A
CIFAR-100
100/100
50,000
10,000
10 × 10
“a good photo of a c .”
A
CUB-200
200/200
5,994
5,794
10 × 20
“a good photo of a c .”
A
ImageNet-R
200/200
24,000
6,000
10 × 20
“a good photo of a c .”
A
ImageNet-100
100/1,000
129,395
5,000
10 × 10
“a good photo of a c .”
B
CIFAR-100
100/100
50,000
10,000
10 × 10
“a photo of a c .”
B
SUN-397
300/397
15,000
15,000
10 × 30
“a photo of a c .”
Table B.1: Dataset statistics and class-incremental protocols. “Classes” reports the number used by the protocol and the number in the original benchmark. “Stream” gives the number of tasks × new classes per task , and “Training template” reports the text prompt used during optimization. Train/test counts refer only to the classes used in the CIL stream.
Method
Additional Data
Setting-A (CLIP)
CIFAR
CUB
I.N.-R
Aˉ
Al
Aˉ
Al
Aˉ
Al
SGCL ( Sharma et al. 2018 )
Replay (20/class)
89.1
82.7
87.1
82.9
86.8
81.9
CLAP4CLIP ( Jha, Gong, and Yao 2024 )
Replay (20/class)
85.1
76.4
85.2
79.9
85.0
79.2
PROOF ( Zhou et al. 2025b )
Replay (20/class)
86.2
76.3
–
–
82.8
77.1
SPU ∗ ( Yu et al. 2024 )
Replay (20/class)
89.2
83.6
81.0
70.3
85.7
80.1
Table C.1: Comparison with methods using additional training data under Setting-A. Replay methods retain 20 historical images per observed class. CC and COCO denote Conceptual Captions and COCO Captions, respectively. SPU ∗ is the replay-based variant, whereas SPU and DuLBE use neither historical exemplars nor external training images. The best results are highlighted in bold , and the second-best results are underlined .
Method
CIFAR
I.N.-100
Aˉ
Al
Aˉ
Al
O-LoRA ( r=8 /task)
86.2
77.3
83.4
72.0
InfLoRA ( r=10 )
87.8
80.9
85.4
75.5
CL-LoRA ( r=10 )
83.9
75.1
83.7
73.4
LoDA ( r=8 )
86.0
80.4
85.6
74.7
LoDA ( r=32 )
86.4
81.1
85.8
75.7
Table C.2: Comparison with LoRA-based methods under Setting-A. All methods use identical insertion locations in the visual tower, including the key and value projections and the MLP layers. All LoRA baselines and DuLBE Text. use the same standard CLIP text classifier, whereas DuLBE B.E. uses the complete reliability-guided bridge-prototype ensemble. For O-LoRA, r=8 denotes the rank of each task-wise adapter. Best and second-best results are shown in bold and underline , respectively.
Method
Setting-B (OpenCLIP)-LFH
CIFAR [-1pt] B50 Inc10
SUN [-1pt] B150 Inc30
Cars [-1pt] B50 Inc10
Food [-1pt] B50 Inc10
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
ContinualCLIP ( Thengane et al. 2022 )
76.5
71.4
75.0
72.1
78.3
76.4
84.8
81.9
DualPrompt ( Wang et al. 2022 )
80.1
72.6
79.4
73.0
76.9
67.6
80.0
72.8
CODA-Prompt ( Smith et al. 2023 )
78.7
71.6
80.4
74.2
75.1
64.2
81.0
74.1
PROOF ( Zhou et al. 2025b )
82.9
78.9
80.7
77.5
90.5
89.5
87.5
84.7
Table C.3: Comparison under the learning-from-half (LFH) protocol in Setting-B. The best results are highlighted in bold , and the second-best results are underlined .
Method
CIFAR
I.N.-R
Aˉ
Al
Aˉ
Al
PROOF ( Zhou et al. 2025b )
89.9
83.6
91.3
87.3
CLAP4CLIP ( Jha, Gong, and Yao 2024 )
87.9
84.9
92.1
88.6
SLCA ( Zhang et al. 2023 )
90.1
84.6
90.0
86.8
RAPF ( Huang et al. 2024 )
90.3
85.3
92.0
88.3
L2P++ ( Luo et al. 2026 )
85.7
77.9
90.5
86.7
Table C.4: Comparison using OpenAI CLIP ViT-L/14 under the uniform Setting-A protocol. CIFAR-100 and ImageNet-R are divided into ten tasks with 10 and 20 new classes per task, respectively. Baseline results follow the matched ViT-L/14 evaluation reported by ( Huang et al. 2025 ) . CLAP4CLIP ∗ denotes its replay-free variant. Text. and B.E. denote the conventional text classifier and bridge-prototype ensemble, respectively. The best results are highlighted in bold , and the second-best results are underlined .
Figure C.1: Effect of fixed bridge depth α and the number of reliability-weighted bridge depths Kb on CIFAR-100 under Setting-A. Blue denotes a fixed-depth classifier, while red denotes the reliability-guided bridge-prototype ensemble.
Figure C.2: Effect of the reliability temperature β with Kb=10 on CIFAR-100 under Setting-A. (a) Distributions of the class-wise expected bridge depth Eπc[α] (top) and reliability weights of three representative classes (bottom). The red vertical line denotes the mean expected depth across classes. (b) Final accuracy under different values of β .
Adapted visual modules
Aˉ↑
Al↑
K,V
88.8
83.0
Q,K,V
89.1
83.9
K,V+MLP (default)
89.5
84.2
MLP
89.2
83.5
Q,K,V+MLP
89.6
84.1
Table C.5: Effect of the insertion locations of the dual-mode low-rank learner on CIFAR-100 under Setting-A. Here, Q , K , and V denote the query, key, and value projections in visual self-attention, while MLP denotes the feed-forward block. The same per-module ranks are used for all configurations. The best results are highlighted in bold , and the second-best results are underlined .
Method
Training Stage
Inference Stage
CIFAR (A)
Trainable Params (M)
GFLOPs (per sample)
Additional Params (M)
Aˉ↑
Al↑
RAPF ( Huang et al. 2024 )
0.26
17.6003
0.262
86.2
79.0
LoRA-CLIP r=32 ( Chen et al. 2015 )
6.88
17.6001
0.000
86.2
79.1
LoDA-CLIP r=32† ( He et al. 2026c )
1.77
35.2001
86.244
86.4
81.1
MG-CLIP ( Huang et al. 2025 )
0.54
17.6001
0.051
87.0
80.6
SGCL ∗ ( Sharma et al. 2018 )
7.09
17.6001
0.000
86.4
79.5
Table C.6: Training and inference overhead on CIFAR-100 under Setting-A. Trainable parameters cover the complete method, including trainable components in both CLIP towers. GFLOPs include visual encoding and classification over all 100 classes. SGCL ∗ denotes the replay-free variant. MG-CLIP uses one image encoding and combines a text classifier with a learned visual classifier. DuLBE directly evaluates Kb=10 bridge depths, incurring Kb∣C∣d=0.000512 GFLOPs for bridge scoring. Additional parameters include method-specific model parameters and persistent prototype buffers retained beyond the standard CLIP model, while excluding the common text-embedding cache. The best results are in bold , and the second-best results are underlined .
Method
Stored state
MiB ↓
Al↑
RAPF
Mean + covariance
100.20
79.0
ENGINE
Prototype + GDA
3.03
73.1
CLG-CBM
Mean + covariance + concepts
102.15
76.9
DuLBE
Prototype + reliability
0.20
84.2
Table C.7: Representation memory and final accuracy on CIFAR-100 under Setting-A.
Al↑
Config.
Learner location
MiB ↓
CIFAR
Food
Full
K,V+MLP
486.00
84.5
86.7
A
K,V , all blocks
27.00
83.9
85.5
B
K,V , last four
9.00
83.6
85.1
C
Bridge layer
2.25
81.2
84.1
BOFA
Bridge layer
2.25
79.3
82.7
Table C.8: Statistic storage and final accuracy under Setting-B.
Figure C.3: Prototype-conditioned Layer Grad × Act maps on ImageNet-100 and ImageNet-R. Columns show the visual prototype, bridge prototypes at α∈{0.25,0.50,0.75} , and the text embedding. Warmer colors indicate larger positive contributions to the image-prototype similarity.
Figure C.4: CIFAR-100 gradient-space diagnostics. (a) Task-layer log expected-gradient energy across visual Transformer layers. (b) Layer-wise expected-gradient capture ratios of the shared and residual spaces, averaged over Task2–Task10 (T2–T10). The two ratios are computed independently and are not complementary.
Class-Incremental Learning (CIL) aims to continually learn new categories without forgetting previously acquired knowledge. Vision-language models such as CLIP offer strong transferable representations via multi-modal supervision, making them promising for CIL. However, applying CLIP to CIL poses two major challenges: (1) adapting to downstream tasks often requires additional learnable modules, increasing model complexity and susceptibility to forgetting; and (2) while multi-modal representations offer complementary strengths, existing methods have yet to fully realize their potential in effectively integrating visual and textual modalities. To address these issues, we propose BOFA (Bridge-layer Orthogonal Fusion for Adaptation), a novel framework for CIL. BOFA confines all model adaptation exclusively to CLIP's existing cross-modal bridge-layer, thereby adding no extra parameters or inference cost. To prevent forgetting within this layer, it leverages Orthogonal Low-Rank Fusion, a mechanism that constrains parameter updates to a low-rank ``safe subspace" mathematically constructed to be orthogonal to past task features. This ensures stable knowledge accumulation without data replay. Furthermore, BOFA employs a cross-modal hybrid prototype that synergizes stable textual prototypes with visual counterparts derived from our stably adapted bridge-layer, enhancing classification performance. Extensive experiments on standard benchmarks show that BOFA achieves superior accuracy and efficiency compared to existing methods.
Lan Li, Tao Hu, Da-Wei Zhou +3
School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China
Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features. Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VIS uses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VIS employs a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VIS accumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VIS achieves state-of-the-art performance without a textual branch.
Tao Hu, Zhen-Hao Xie, Jingcai Guo +2
School of Artificial Intelligence, Nanjing University · State Key Laboratory for Novel Software Technology, Nanjing University · Hong Kong Polytechnic University
Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tasks, and without access to task identity at inference. Parameter-efficient methods built on frozen vision-language models have made strong progress, yet all existing approaches rely exclusively on visual features for task routing, confidence estimation, and encoder adaptation, leaving CLIP's cross-modal text embedding space entirely unexploited. We address this gap through three contributions. Text-space task routing replaces visual Gaussian matching with cosine similarity to frozen CLIP text prototypes, giving order-independent routing robust to data scarcity at zero parameter cost. Multi-prototype visual-textual confidence replaces single-Gaussian class modeling with K-means visual prototypes and cross-modal alignment scores under task-calibrated thresholds. Symmetric cross-modal gating extends per-layer Gumbel gates to the text encoder conditioned on batch image features, preserving cross-modal alignment on out-of-distribution inputs. On the MTIL benchmark spanning 11 datasets and 1201 classes, our method achieves 74.2% Transfer, 80.5% Average, and 88.7% Last under Order-I, surpassing the prior state of the art by 5.0, 3.7, and 3.0 percentage points with only 2.5M trainable parameters and no external data.