Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.
Figures & tables
Figure 1: Motivating observations. (a) New-task cross-modal gradients retain substantial projection energy in historically occupied directions ( Kh=128 ) across the visual tower. (b) Tuning occupied directions causes greater old-class relation-structure drift than tuning low-interference directions. SR denotes rank-structure similarity to the old-task relation matrix, whose entries are cosine similarities between old-class visual prototypes and corresponding text embeddings.
Figure 2: Overview of DuLBE . (a) Dual-mode allocation uses the prospective task gradient to form a compact shared mode within historically occupied directions and a low-interference residual mode outside them. (b) Routed cross-modal training directs gradients from the current-task objective to both learners, while routing gradients from the semantic-structure preservation objective only to the shared learner. (c) Classifier: it builds visual-text geodesic bridges, estimates bridge-depth reliability after each task, and ensembles reliable bridge prototypes for inference.
Method
Setting-A (CLIP)
Setting-B (OpenCLIP)
CIFAR
CUB
I.N.-R
I.N.100
CIFAR
SUN
Cars
Food
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
ContinualCLIP
–
66.7
–
51.2
–
72.0
–
75.4
–
71.4
–
72.1
–
76.4
–
81.9
DualPrompt (ECCV’22)
81.5
72.5
–
–
82.0
75.8
80.7
67.4
81.6
72.4
82.5
74.4
76.3
62.9
84.9
77.3
CODA (CVPR’23)
77.0
62.3
66.6
50.9
78.0
67.5
64.1
34.8
82.4
73.4
83.3
75.7
80.2
66.5
86.2
78.8
PROOF (TPAMI’25)
86.2
76.3
–
–
82.8
77.1
84.7
72.5
86.8
79.1
83.9
77.3
90.7
86.5
90.0
84.7
Table 1: Main comparison under two standard CLIP-based CIL settings. Setting-A ( Huang et al. 2024 ) uses OpenAI CLIP ViT-B/16, and Setting-B ( Zhou et al. 2025a ) uses OpenCLIP ViT-B/16 pretrained on LAION-400M. All methods are evaluated using the standard 10-task split without external training data . The best results are highlighted in bold , and the second-best results are underlined .
Mode bases
∇Lc-ce
∇Lbi-kl
CIFAR (A)
CUB (A)
BS
BR
BS
BR
Aˉ↑
Al↑
Aˉ↑
Al↑
LoRA ∗ ( r=32 )
–
–
–
–
(86.1)
(78.5)
(81.0)
(73.8)
Shared PS
✓
–
✓
×
+2.4
+4.1
+3.2
+4.0
Residual PR
–
✓
×
×
+2.5
+4.4
+4.7
+5.5
Dual [PS,PR]
✓
✓
×
×
+1.3
+2.4
+1.1
+1.5
Dual [PS,PR]
✓
✓
✓
✓
+2.6
+4.4
+4.2
+5.3
Table 2: Ablation of dual-mode collaborative learning. All values are improvements over the LoRA ∗ baseline, which replaces the low-rank learner with standard LoRA while keeping all other components unchanged.
Method
Text Clf.
B.E. Clf.
CUB
Food
α
Rel.
CUB
Food
Ours
67.9
83.7
0.25
×
77.9
85.0
Ours
–
–
0.50
×
77.7
83.9
Ours
–
–
0.75
×
72.9
84.2
Ours
–
–
{αk}
×
79.1
84.6
Ours
–
–
{αk}
✓
80.6 (+12.7)
86.7 (+3.0)
Table 3: Ablation of the bridge-prototype ensemble on CUB (A) and Food (B). All entries report the last accuracy Al . B.E. denotes bridge-prototype ensemble and Rel. denotes reliability weighting over bridge depths.
Setting
Dataset
Classes
Train
Test
Stream
Training template
A
CIFAR-100
100/100
50,000
10,000
10 × 10
“a good photo of a c .”
A
CUB-200
200/200
5,994
5,794
10 × 20
“a good photo of a c .”
A
ImageNet-R
200/200
24,000
6,000
10 × 20
“a good photo of a c .”
A
ImageNet-100
100/1,000
129,395
5,000
10 × 10
“a good photo of a c .”
B
CIFAR-100
100/100
50,000
10,000
10 × 10
“a photo of a c .”
B
SUN-397
300/397
15,000
15,000
10 × 30
“a photo of a c .”
Table B.1: Dataset statistics and class-incremental protocols. “Classes” reports the number used by the protocol and the number in the original benchmark. “Stream” gives the number of tasks × new classes per task , and “Training template” reports the text prompt used during optimization. Train/test counts refer only to the classes used in the CIL stream.
Method
Additional Data
Setting-A (CLIP)
CIFAR
CUB
I.N.-R
Aˉ
Al
Aˉ
Al
Aˉ
Al
SGCL ( Sharma et al. 2018 )
Replay (20/class)
89.1
82.7
87.1
82.9
86.8
81.9
CLAP4CLIP ( Jha, Gong, and Yao 2024 )
Replay (20/class)
85.1
76.4
85.2
79.9
85.0
79.2
PROOF ( Zhou et al. 2025b )
Replay (20/class)
86.2
76.3
–
–
82.8
77.1
SPU ∗ ( Yu et al. 2024 )
Replay (20/class)
89.2
83.6
81.0
70.3
85.7
80.1
Table C.1: Comparison with methods using additional training data under Setting-A. Replay methods retain 20 historical images per observed class. CC and COCO denote Conceptual Captions and COCO Captions, respectively. SPU ∗ is the replay-based variant, whereas SPU and DuLBE use neither historical exemplars nor external training images. The best results are highlighted in bold , and the second-best results are underlined .
Method
CIFAR
I.N.-100
Aˉ
Al
Aˉ
Al
O-LoRA ( r=8 /task)
86.2
77.3
83.4
72.0
InfLoRA ( r=10 )
87.8
80.9
85.4
75.5
CL-LoRA ( r=10 )
83.9
75.1
83.7
73.4
LoDA ( r=8 )
86.0
80.4
85.6
74.7
LoDA ( r=32 )
86.4
81.1
85.8
75.7
Table C.2: Comparison with LoRA-based methods under Setting-A. All methods use identical insertion locations in the visual tower, including the key and value projections and the MLP layers. All LoRA baselines and DuLBE Text. use the same standard CLIP text classifier, whereas DuLBE B.E. uses the complete reliability-guided bridge-prototype ensemble. For O-LoRA, r=8 denotes the rank of each task-wise adapter. Best and second-best results are shown in bold and underline , respectively.
Method
Setting-B (OpenCLIP)-LFH
CIFAR [-1pt] B50 Inc10
SUN [-1pt] B150 Inc30
Cars [-1pt] B50 Inc10
Food [-1pt] B50 Inc10
Aˉ
Al
Aˉ
Al
Aˉ
Al
Aˉ
Al
ContinualCLIP ( Thengane et al. 2022 )
76.5
71.4
75.0
72.1
78.3
76.4
84.8
81.9
DualPrompt ( Wang et al. 2022 )
80.1
72.6
79.4
73.0
76.9
67.6
80.0
72.8
CODA-Prompt ( Smith et al. 2023 )
78.7
71.6
80.4
74.2
75.1
64.2
81.0
74.1
PROOF ( Zhou et al. 2025b )
82.9
78.9
80.7
77.5
90.5
89.5
87.5
84.7
Table C.3: Comparison under the learning-from-half (LFH) protocol in Setting-B. The best results are highlighted in bold , and the second-best results are underlined .
Method
CIFAR
I.N.-R
Aˉ
Al
Aˉ
Al
PROOF ( Zhou et al. 2025b )
89.9
83.6
91.3
87.3
CLAP4CLIP ( Jha, Gong, and Yao 2024 )
87.9
84.9
92.1
88.6
SLCA ( Zhang et al. 2023 )
90.1
84.6
90.0
86.8
RAPF ( Huang et al. 2024 )
90.3
85.3
92.0
88.3
L2P++ ( Luo et al. 2026 )
85.7
77.9
90.5
86.7
Table C.4: Comparison using OpenAI CLIP ViT-L/14 under the uniform Setting-A protocol. CIFAR-100 and ImageNet-R are divided into ten tasks with 10 and 20 new classes per task, respectively. Baseline results follow the matched ViT-L/14 evaluation reported by ( Huang et al. 2025 ) . CLAP4CLIP ∗ denotes its replay-free variant. Text. and B.E. denote the conventional text classifier and bridge-prototype ensemble, respectively. The best results are highlighted in bold , and the second-best results are underlined .
Figure C.1: Effect of fixed bridge depth α and the number of reliability-weighted bridge depths Kb on CIFAR-100 under Setting-A. Blue denotes a fixed-depth classifier, while red denotes the reliability-guided bridge-prototype ensemble.
Figure C.2: Effect of the reliability temperature β with Kb=10 on CIFAR-100 under Setting-A. (a) Distributions of the class-wise expected bridge depth Eπc[α] (top) and reliability weights of three representative classes (bottom). The red vertical line denotes the mean expected depth across classes. (b) Final accuracy under different values of β .
Adapted visual modules
Aˉ↑
Al↑
K,V
88.8
83.0
Q,K,V
89.1
83.9
K,V+MLP (default)
89.5
84.2
MLP
89.2
83.5
Q,K,V+MLP
89.6
84.1
Table C.5: Effect of the insertion locations of the dual-mode low-rank learner on CIFAR-100 under Setting-A. Here, Q , K , and V denote the query, key, and value projections in visual self-attention, while MLP denotes the feed-forward block. The same per-module ranks are used for all configurations. The best results are highlighted in bold , and the second-best results are underlined .
Method
Training Stage
Inference Stage
CIFAR (A)
Trainable Params (M)
GFLOPs (per sample)
Additional Params (M)
Aˉ↑
Al↑
RAPF ( Huang et al. 2024 )
0.26
17.6003
0.262
86.2
79.0
LoRA-CLIP r=32 ( Chen et al. 2015 )
6.88
17.6001
0.000
86.2
79.1
LoDA-CLIP r=32† ( He et al. 2026c )
1.77
35.2001
86.244
86.4
81.1
MG-CLIP ( Huang et al. 2025 )
0.54
17.6001
0.051
87.0
80.6
SGCL ∗ ( Sharma et al. 2018 )
7.09
17.6001
0.000
86.4
79.5
Table C.6: Training and inference overhead on CIFAR-100 under Setting-A. Trainable parameters cover the complete method, including trainable components in both CLIP towers. GFLOPs include visual encoding and classification over all 100 classes. SGCL ∗ denotes the replay-free variant. MG-CLIP uses one image encoding and combines a text classifier with a learned visual classifier. DuLBE directly evaluates Kb=10 bridge depths, incurring Kb∣C∣d=0.000512 GFLOPs for bridge scoring. Additional parameters include method-specific model parameters and persistent prototype buffers retained beyond the standard CLIP model, while excluding the common text-embedding cache. The best results are in bold , and the second-best results are underlined .
Method
Stored state
MiB ↓
Al↑
RAPF
Mean + covariance
100.20
79.0
ENGINE
Prototype + GDA
3.03
73.1
CLG-CBM
Mean + covariance + concepts
102.15
76.9
DuLBE
Prototype + reliability
0.20
84.2
Table C.7: Representation memory and final accuracy on CIFAR-100 under Setting-A.
Al↑
Config.
Learner location
MiB ↓
CIFAR
Food
Full
K,V+MLP
486.00
84.5
86.7
A
K,V , all blocks
27.00
83.9
85.5
B
K,V , last four
9.00
83.6
85.1
C
Bridge layer
2.25
81.2
84.1
BOFA
Bridge layer
2.25
79.3
82.7
Table C.8: Statistic storage and final accuracy under Setting-B.
Figure C.3: Prototype-conditioned Layer Grad × Act maps on ImageNet-100 and ImageNet-R. Columns show the visual prototype, bridge prototypes at α∈{0.25,0.50,0.75} , and the text embedding. Warmer colors indicate larger positive contributions to the image-prototype similarity.
Figure C.4: CIFAR-100 gradient-space diagnostics. (a) Task-layer log expected-gradient energy across visual Transformer layers. (b) Layer-wise expected-gradient capture ratios of the shared and residual spaces, averaged over Task2–Task10 (T2–T10). The two ratios are computed independently and are not complementary.
School of Artificial Intelligence, Nanjing University · State Key Laboratory for Novel Software Technology, Nanjing University · Hong Kong Polytechnic University