Organizations: Department of Mathematics, Alma Mater Studiorum – Università di Bologna, Piazza di Porta San Donato 5, 40126 Bologna, Italy · Gatsby Computational Neuroscience Unit, University College London · Donders Centre for Neuroscience, Radboud University, Nijmegen, The Netherlands
Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.
Figures & tables
Figure 1: Generalization and representation dynamics under LoRA and full fine-tuning for a sequential training. a) Schematic of the sequential training setting. We first train the first-layer weights and a readout head on a synthetic dataset corresponding to Task 1. We then train on a correlated dataset corresponding to Task 2, where we either update the first-layer weights or freeze them and instead train a low-rank adapter. In both cases, we train a second readout head. b) Task 1 (dark) and Task 2 (light) generalization errors for full fine-tuning and LoRA. The gray dashed line marks the task switch. c–d) Student overlaps with teachers as defined in Eqs. ( 18 )-( 19 ) after the task switch for both full fine-tuning (red) and LoRA (blue). Solid lines denote theoretical ODE predictions and markers finite-dimensional simulations. Parameters: N=103 , K=10 , M=5 , L=5 , c=0.5 , α=50 .
Figure 2: State-dependent masking reduces forgetting. a) Generalization dynamics for full fine-tuning, LoRA, and their SDGM-constrained variants. Parameters: N=103 , K=10 , M=5 , L=5 , κ=5 , c=0.5 , α=50 . b) Corresponding results on a real experiment on the MNIST dataset sequential training (see Appendix G for details). SDGM improves Task 1 retention while preserving comparable Task 2 performance. For the experiment, the curves are averaged over 10 independent training realizations.
Figure 3: Impact of LoRA on the Stability-Plasticity Trade-off. a) Task 1 Forgetting and b) Task 2 Transfer as a function of teacher similarity c . Solid lines denote theoretical predictions and markers simulation. c) Forgetting versus Transfer as the LoRA rank L varies from light ( L=1 ) to dark marker ( L=10 ) for c=0.5 . Vanilla LoRA gains transfer at the cost of increased forgetting, whereas LoRA+SDGM maintains greater stability. Transfer gains saturate as L approaches the teacher width M . For SDGM, κ=K−L . We average over 10 different seeds. Lines in panel c are shown to guide the eyes. Parameters: N=103 , K=10 , M=5 , α=50 . In panel a and b we used L=5 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Ablation of the selection protocol via Inverse SDGM. Generalization error trajectories on Task 1 and Task 2 under the inverse selection protocol, where student hidden units corresponding to the smallest Task 1 readout magnitudes are frozen during Task 2 training. Comparing this control to standard SDGM disentangles the effect of purely architectural capacity constraints from targeted feature protection. Parameters: N=103K=10 , M=5 , L=5 , c=0.5 , α=50 , κ=5 .
Figure 5: Validation of SDGM under unbounded ReLU activations. Generalization error dynamics on Task 1 and Task 2 when both teacher and student networks employ ReLU as the activation function. The lower error on Task 1 confirms that the proposed selection rule mitigates catastrophic forgetting independently of activation saturation. Parameters: N=103 , K=10 , M=5 , L=5 , c=0.5 , α=50 . Results shown here are experiment-only.
Figure 6: Typical forgetting on the first task and transfer on the second task with the new initialization . In this setting, the SDGM procedure allows for no forgetting on the full range of task similarity, while allowing for the same transfer. The smaller transfer at big task similarity for LoRA + SDGM is due to long symmetric plateau, which size increases non-monotonically with c . Parameters : N= 103 , K= 10 , M= 5 , L= 5 , α=50 .
Figure 7: Additional results in the specialized regimes. a) Typical generalization error for full fine-tuning (red), LoRA (blue), and their SDGM-constrained variants. During Task 1 training, the symmetric plateau, corresponding to the absence of specialization, is located at logϵ∗≈−1.5 . b) Task 1 forgetting and c) Task 2 transfer as a function of teacher similarity c . d–e) Student overlaps with the teachers, as defined in Eqs.( 18 )–( 19 ), after the task switch for LoRA + SDGM (green) and standard + SDGM (yellow). The overlaps corresponding to frozen student directions remain constant throughout Task 2 training.
Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning (CL), the standard assumption is that each new task requires a dedicated low-rank adapter. In this work, we challenge this assumption empirically and structurally. We show that task-specific LoRA adapters in CL exhibit significant low-rank redundancy: the subspaces spanned by adapters trained on different tasks substantially overlap, and in many cases earlier adapters can faithfully represent later tasks. Building on this observation, we propose LiteLoRA, a plug-and-play gating mechanism that learns at train time whether to recruit a new adapter or reuse existing low-rank representations. Our method reduces the number of active adapters by 20-70% while matching or exceeding state-of-the-art performance on standard CL benchmarks, revealing that structural redundancy is pervasive and that selective learning is sufficient to achieve stability without sacrificing plasticity.
While orthogonal subspace methods try to mitigate task interference in Continual Learning (CL), they often suffer from energy diffusion across the basis, hindering knowledge compaction and exhausting capacity for future tasks. We observe that output feature drift induced by parameter updates is inherently low-rank, and theoretically prove that preserving parameters along the principal directions of this drift minimizes the output reconstruction error. Motivated by this, we propose \textbf{E}nergy-Concentrated and \textbf{E}nergy-Ordered \textbf{Lo}w-\textbf{R}ank \textbf{A}daptation (E2-LoRA). By explicitly ordering and concentrating knowledge into leading ranks, E2-LoRA frees capacity for subsequent tasks. Furthermore, we design a dynamic rank allocation strategy to balance stability and plasticity by jointly optimizing energy retention and model plasticity. Extensive experiments across multiple benchmarks demonstrate that E2-LoRA achieves state-of-the-art performance. Code is available at https://github.com/kiddo127/E2-LoRA.
Longhua Li, Lei Qi, Qi Tian +1
School of Computer Science and Engineering, Southeast University, Nanjing, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China · Huawei Technologies, Shenzhen, China
Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge. Most existing approaches treat continual learning as avoiding interference with past updates, rather than considering what properties make the current task-specific update naturally preserve previously acquired knowledge. From a knowledge-decomposition perspective, we observe that low-rank adaptations exhibit highly imbalanced singular value spectra: a few dominant components absorb most of the adaptation energy, thereby (i) more likely to disrupt previously acquired knowledge and (ii) making the update more vulnerable to interference from subsequent tasks. To enable explicit balance among components, we decouple the magnitude of the task update from its directional structure and formulate it as a constrained optimization problem on a restricted Stiefel manifold. We address this problem using a projected first-order method compatible with standard deep-learning optimizers used in vision-language models. Our method mitigates both backward and forward forgetting, consistently outperforming continual learning baselines. The implementation code is available at https://github.com/haodotgu/EBLoRA.
Hao Gu, Mao-Lin Luo, Zi-Hao Zhou +3
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China