C-LoRA: Continual Low-Rank Adaptation for Pre-trained Visual Models
Authors: Xin Zhang, Liang Bai, Xian Yang
Organizations: Institute of Intelligent Information Processing, Shanxi University, Taiyuan, 030006, China · Alliance Manchester Business School, The University of Manchester, Manchester, M13 9PL, UK
Pre-trained visual models have become fundamental in computer vision, but they face challenges in continual learning scenarios where data and tasks evolve over time. Low-Rank Adaptation (LoRA) offers efficient fine-tuning capabilities but remains limited for such dynamic environments. Standard LoRA cannot distinguish important subspaces, causing critical knowledge to be overwritten in sequential training. Existing approaches address this by dynamically expanding the set of LoRA adapters, either maintaining a growing pool of task-specific modules or merging new adapters into prior ones, at the cost of unbounded parameter growth or increasing inference complexity. We propose Continual Low-Rank Adaptation (C-LoRA), a method that enables a single, shared LoRA adapter to handle sequential tasks without catastrophic forgetting, without requiring any module selection or fusion at inference. The core of C-LoRA is a learnable routing matrix R that explicitly controls how each rank-one subspace contributes to the weight update. This matrix is decomposed into a stability component (R_base), which preserves knowledge from prior tasks, and a plasticity component (R_delta), which drives adaptation to the current task, providing direct control over the stability-plasticity trade-off. We analyze how R governs gradient flow during sequential training, and demonstrate competitive performance across multiple benchmarks.
Figures & tables
Figure 1: (a) The standard LoRA formulation implicitly assumes fixed, uniform weights (implicitly 1) and only allows interaction between corresponding rank-one subspaces ( aibi⊤ ). (b) C-LoRA introduces a learnable routing matrix R , enabling fine-grained control over subspace contributions and capturing complex cross-subspace interactions between ai and bj . (c) The inherent conflict in continual learning: High-magnitude areas in R (hotspots) are crucial for stability (preserving past knowledge), but they simultaneously amplify the gradients flowing to the shared matrices A and B during backpropagation, thereby risking catastrophic forgetting.
Figure 2: (a): Heatmap of the R matrix at initialization, showing a nearly uniform distribution; (b): Heatmap of the R matrix after training, with noticeable localised value changes; (c): Variance of matrix elements increasing with training epochs; (d): Range of matrix values expanding as training progresses.
Figure 3: Evolution of the routing matrix R heatmap during 5-task class-incremental learning. Top Row (Before Decomposition): Without decomposition, R values progressively intensify in specific local regions. This highlights that subspaces critical for past tasks continue to receive strong updates when learning new tasks. Bottom Row (After Decomposition): With decomposition, updates primarily occur in the task-specific Rδ , resulting in changes to R that are more distributed across the matrix rather than being highly localized.
Figure 4: Illustration of the Proposed Model. Left: Vision Transformer (ViT) integrated with the C-LoRA module, where the adapter and local classifier are incrementally trained in each session. Right: Proposed architecture mitigates catastrophic forgetting by decoupling R . This isolates updates within a task-specific component, Rδ , which in turn drives distributed updates to discover new routing configurations while preserving previously learned knowledge.
Method
Split CIFAR-100
Split ImageNet-A
Split CUB-200
Split CAR196
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Joint-Training
93.22
-
79.61
-
90.78
-
78.66
-
EWC Kirkpatrick et al. (2017)
90.59 ± 0.63
93.80 ± 0.97
56.24 ± 1.56
66.30 ± 1.71
85.10 ± 0.31
92.36 ± 0.51
45.58 ± 0.19
60.41 ± 0.94
L2P Wang et al. (2022c)
84.17 ± 0.91
88.20 ± 1.30
43.47 ± 0.89
50.52 ± 2.29
67.13 ± 2.03
79.62 ± 1.56
45.81 ± 0.98
57.85 ± 1.92
DualPrompt Wang et al. (2022b)
81.70 ± 1.07
86.42 ± 0.98
46.41 ± 0.20
55.95 ± 1.34
68.68 ± 0.35
80.75 ± 0.87
39.36 ± 2.32
53.93 ± 2.40
SLCA Zhang et al. (2023a)
91.07 ± 0.45
94.06 ± 1.07
58.96 ± 0.88
66.39 ± 2.16
84.55 ± 0.19
90.69 ± 0.65
66.49 ± 2.87
76.43 ± 0.61
Table 1: Accuracy comparison of different continual learning methods across four benchmark datasets, divided into 10 incremental sessions.
Figure 5: Accuracy performance during the 10 incremental sessions on CIFAR-100, ImageNet-A, CUB-200, and CAR196.
Method
Total Expanded Parameters (M)
Trainable Parameters per Task (M)
EASE
23.6
1.18
InfLoRA
7.4
0.37
SD-LoRA
7.4
0.37
C-LoRA (Ours)
0.36
0.36
Table 2: Comparison of parameter efficiency for different LoRA-based continual learning methods under a 20-task setting. C-LoRA maintains a fixed parameter budget, while others grow linearly with the number of tasks.
Method
ImageNet-A
CUB-200
Last-Acc (%)
Inc-Acc (%)
BWT
Last-Acc (%)
Inc-Acc (%)
BWT
w/o Decomp & Gate
55.76
67.15
−28.90
85.67
90.61
−11.78
w/o Gate
62.21
71.09
−13.33
89.74
93.36
−4.47
C-LoRA (Ours)
64.06
73.38
−9.77
90.25
93.63
−3.78
Table 3: Ablation study on the core components of C-LoRA (10 sessions).
Figure 6: Per-task accuracy on ImageNet-A (left) and CUB-200 (right).
Method
Split CIFAR-100
Split ImageNet-R
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
LoRA-Expansion Methods
EASE
87.53 ± 0.38
91.67 ± 1.04
74.52 ± 0.93
80.22 ± 1.31
InfLoRA
84.31 ± 0.28
90.31 ± 0.28
74.75 ± 0.64
80.67 ± 0.55
SD-LoRA
86.65 ± 1.54
90.78 ± 1.25
75.72 ± 1.15
80.39 ± 0.82
C-LoRA (Ours)
Table 4: Comparison with LoRA-expansion methods and ablation on the feature resampling strategy (Stage 2) on CIFAR-100 and ImageNet-R (10 sessions).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Split CIFAR-100
Split ImageNet-A
Split CUB-200
Split CAR196
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Joint-Training
93.22
-
79.61
-
90.78
-
78.66
-
EWC Kirkpatrick et al. [2017]
91.21 ± 0.53
93.93 ± 0.67
60.13 ± 2.28
68.74 ± 0.87
84.73 ± 0.74
91.07 ± 0.38
51.64 ± 1.36
61.47 ± 1.10
L2P Wang et al. [2022c]
84.60 ± 1.74
89.24 ± 1.27
46.94 ± 0.71
53.28 ± 1.13
73.07 ± 1.74
81.99 ± 1.34
56.03 ± 2.86
66.43 ± 2.40
DualPrompt Wang et al. [2022b]
82.66 ± 2.00
87.23 ± 0.88
50.54 ± 0.98
58.18 ± 1.30
74.28 ± 0.26
83.60 ± 0.44
48.90 ± 0.76
61.83 ± 0.44
SLCA Zhang et al. [2023a]
92.19 ± 0.56
94.56 ± 0.58
63.14 ± 1.39
68.65 ± 0.72
86.54 ± 1.27
91.51 ± 0.96
75.73 ± 1.27
82.01 ± 0.72
Appendix
Table 5: Accuracy comparison of different continual learning methods across four benchmark datasets, divided into 5 sessions.
Figure 7: Accuracy performance during the 5 sessions on CIFAR-100, ImageNet-A, CUB-200, and CAR196.
Method
Split CIFAR-100
Split ImageNet-A
Split CUB-200
Split CAR196
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
Joint-Training
93.22
-
79.61
-
90.78
-
78.66
-
EWC Kirkpatrick et al. [2017]
89.95 ± 0.15
93.31 ± 1.02
54.62 ± 1.39
63.55 ± 1.53
87.25 ± 1.00
93.45 ± 0.48
44.42 ± 0.56
61.75 ± 1.06
L2P Wang et al. [2022c]
79.32 ± 1.49
84.10 ± 1.11
40.07 ± 1.42
49.35 ± 1.37
60.30 ± 3.80
73.53 ± 3.34
29.83 ± 3.31
42.82 ± 1.14
DualPrompt Wang et al. [2022b]
76.57 ± 0.60
82.97 ± 1.78
41.48 ± 1.63
52.84 ± 1.57
65.22 ± 1.29
77.86 ± 1.31
25.11 ± 2.09
40.89 ± 0.77
SLCA Zhang et al. [2023a]
90.37 ± 0.56
93.64 ± 0.99
53.76 ± 5.54
61.82 ± 4.74
82.37 ± 0.42
90.12 ± 0.93
56.09 ± 3.89
68.50 ± 2.42
Appendix
Table 6: Accuracy comparison of different continual learning methods across four benchmark datasets, divided into 20 sessions. This setting emphasizes the challenge of longer task sequences.
Figure 8: Accuracy performance during the 20 sessions on CIFAR-100, ImageNet-A, CUB-200, and CAR196.
Method
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
Avg
w/o Decomp & Gate
63.2
75.4
83.9
85.8
86.2
90.3
88.4
93.2
93.7
94.8
85.5
w/o Gate
79.8
89.4
89.0
89.5
92.2
93.4
89.3
93.6
92.0
89.6
89.8
C-LoRA (Ours)
83.8
90.8
90.4
89.1
92.2
92.5
90.2
92.8
91.6
89.1
90.3
Appendix
Table 7: CUB-200 per-task accuracy (%) after learning all tasks.
Method
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
Avg
w/o Decomp & Gate
36.0
21.7
11.5
10.9
10.4
4.8
5.3
4.0
1.3
0.0
10.6
w/o Gate
18.6
8.2
5.1
4.9
2.6
−0.4
0.5
−0.8
1.7
0.0
4.0
C-LoRA (Ours)
14.6
6.8
4.1
6.1
2.2
−0.9
−1.3
1.6
0.8
0.0
3.4
Appendix
Table 8: CUB-200 per-task forgetting (%) after learning all tasks.
Figure 9: CUB-200: Per-Task Final Accuracy & Forgetting Reduction
Session
∥∂L/∂A∥F
∥R∥2
C-LoRA
w/o Decomp
Ratio
C-LoRA
w/o Decomp
Ratio
ImageNet-A
1
0.0013
0.0506
39 ×
1.09
1.56
1.43 ×
5
0.0005
0.1118
234 ×
1.17
2.65
2.26 ×
9
0.0007
0.1638
236 ×
1.24
3.25
2.63 ×
CUB-200
Appendix
Table 9: Gradient norm ∥∂L/∂A∥F and routing-matrix norm ∥R∥2 across sessions (12-layer mean, 10 sessions). “w/o Decomp” denotes the variant without decomposition and gating; Ratio = w/o Decomp / C-LoRA.
Method
ImageNet-A
CUB-200
Cars-196
Last-Acc
Inc-Acc
Last-Acc
Inc-Acc
Last-Acc
Inc-Acc
Vanilla LoRA
53.85
–
85.33
–
69.38
–
w/o Decomp & Gate
55.76
67.15
85.67
90.61
69.69
79.19
+ LR scaling
60.11
69.85
89.99
–
70.44
78.98
C-LoRA (Ours)
64.06
73.38
90.25
93.63
71.97
80.01
Appendix
Table 10: Controlled comparison of capacity and learning-rate scaling (10 sessions). “–” denotes not evaluated.
Method
ImageNet-A
CUB-200
Last-Acc (%)
Inc-Acc (%)
Last-Acc (%)
Inc-Acc (%)
AdaLoRA
57.41
68.56
83.67
91.52
CL-LoRA
58.92
70.06
80.03
87.66
C-LoRA (Ours)
64.06
73.38
90.25
93.63
Appendix
Table 11: Comparison with AdaLoRA and CL-LoRA on ImageNet-A and CUB-200 (10 sessions).
Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning (CL), the standard assumption is that each new task requires a dedicated low-rank adapter. In this work, we challenge this assumption empirically and structurally. We show that task-specific LoRA adapters in CL exhibit significant low-rank redundancy: the subspaces spanned by adapters trained on different tasks substantially overlap, and in many cases earlier adapters can faithfully represent later tasks. Building on this observation, we propose LiteLoRA, a plug-and-play gating mechanism that learns at train time whether to recruit a new adapter or reuse existing low-rank representations. Our method reduces the number of active adapters by 20-70% while matching or exceeding state-of-the-art performance on standard CL benchmarks, revealing that structural redundancy is pervasive and that selective learning is sufficient to achieve stability without sacrificing plasticity.
Low-Rank Adaptation (LoRA) has emerged as a promising paradigm for Continual Learning. It independently updates its low-rank factors (A and B), creating a composite update to the full weight matrix through their interaction. To prevent catastrophic forgetting, this update should remain orthogonal to the task-specific subspace that contains previously learned knowledge. However, we identify that this composite update systematically violates this orthogonality, reintroducing interference and undermining stability. Furthermore, naively enforcing this orthogonality compromises plasticity, disrupting the delicate stability-plasticity trade-off. To resolve these issues, we propose \textbf{Janus-LoRA}, a framework that restores this balance through two novel components. Specifically, we first introduce Gradient Rectification, a closed-form solution that mathematically decouples LoRA's factor updates, enforcing orthogonality against the historical knowledge subspace identified by an efficient Online Estimation. Next, to enhance plasticity, we introduce a Decoupled Margin Loss that promotes feature-level separation by pushing new feature representations away from old ones, thus creating distinct, low-interference regions for new learning. Comprehensive experiments on challenging benchmarks demonstrate that by harmonizing parameter-level orthogonality with feature-level separation, Janus-LoRA achieves a superior balance and establishes new state-of-the-art performance.
Cheng Chen, Pengpeng Zeng, Yuyu Guo +3
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, China · School of Computer Science and Technology, Tongji University, Shanghai, China · Independent Researcher +1
While orthogonal subspace methods try to mitigate task interference in Continual Learning (CL), they often suffer from energy diffusion across the basis, hindering knowledge compaction and exhausting capacity for future tasks. We observe that output feature drift induced by parameter updates is inherently low-rank, and theoretically prove that preserving parameters along the principal directions of this drift minimizes the output reconstruction error. Motivated by this, we propose \textbf{E}nergy-Concentrated and \textbf{E}nergy-Ordered \textbf{Lo}w-\textbf{R}ank \textbf{A}daptation (E2-LoRA). By explicitly ordering and concentrating knowledge into leading ranks, E2-LoRA frees capacity for subsequent tasks. Furthermore, we design a dynamic rank allocation strategy to balance stability and plasticity by jointly optimizing energy retention and model plasticity. Extensive experiments across multiple benchmarks demonstrate that E2-LoRA achieves state-of-the-art performance. Code is available at https://github.com/kiddo127/E2-LoRA.
Longhua Li, Lei Qi, Qi Tian +1
School of Computer Science and Engineering, Southeast University, Nanjing, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China · Huawei Technologies, Shenzhen, China