Organizations: Department of Mathematics, Alma Mater Studiorum – Università di Bologna, Piazza di Porta San Donato 5, 40126 Bologna, Italy · Gatsby Computational Neuroscience Unit, University College London · Donders Centre for Neuroscience, Radboud University, Nijmegen, The Netherlands
Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.
Figures & tables
Figure 1: Generalization and representation dynamics under LoRA and full fine-tuning for a sequential training. a) Schematic of the sequential training setting. We first train the first-layer weights and a readout head on a synthetic dataset corresponding to Task 1. We then train on a correlated dataset corresponding to Task 2, where we either update the first-layer weights or freeze them and instead train a low-rank adapter. In both cases, we train a second readout head. b) Task 1 (dark) and Task 2 (light) generalization errors for full fine-tuning and LoRA. The gray dashed line marks the task switch. c–d) Student overlaps with teachers as defined in Eqs. ( 18 )-( 19 ) after the task switch for both full fine-tuning (red) and LoRA (blue). Solid lines denote theoretical ODE predictions and markers finite-dimensional simulations. Parameters: N=103 , K=10 , M=5 , L=5 , c=0.5 , α=50 .
Figure 2: State-dependent masking reduces forgetting. a) Generalization dynamics for full fine-tuning, LoRA, and their SDGM-constrained variants. Parameters: N=103 , K=10 , M=5 , L=5 , κ=5 , c=0.5 , α=50 . b) Corresponding results on a real experiment on the MNIST dataset sequential training (see Appendix G for details). SDGM improves Task 1 retention while preserving comparable Task 2 performance. For the experiment, the curves are averaged over 10 independent training realizations.
Figure 3: Impact of LoRA on the Stability-Plasticity Trade-off. a) Task 1 Forgetting and b) Task 2 Transfer as a function of teacher similarity c . Solid lines denote theoretical predictions and markers simulation. c) Forgetting versus Transfer as the LoRA rank L varies from light ( L=1 ) to dark marker ( L=10 ) for c=0.5 . Vanilla LoRA gains transfer at the cost of increased forgetting, whereas LoRA+SDGM maintains greater stability. Transfer gains saturate as L approaches the teacher width M . For SDGM, κ=K−L . We average over 10 different seeds. Lines in panel c are shown to guide the eyes. Parameters: N=103 , K=10 , M=5 , α=50 . In panel a and b we used L=5 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Ablation of the selection protocol via Inverse SDGM. Generalization error trajectories on Task 1 and Task 2 under the inverse selection protocol, where student hidden units corresponding to the smallest Task 1 readout magnitudes are frozen during Task 2 training. Comparing this control to standard SDGM disentangles the effect of purely architectural capacity constraints from targeted feature protection. Parameters: N=103K=10 , M=5 , L=5 , c=0.5 , α=50 , κ=5 .
Figure 5: Validation of SDGM under unbounded ReLU activations. Generalization error dynamics on Task 1 and Task 2 when both teacher and student networks employ ReLU as the activation function. The lower error on Task 1 confirms that the proposed selection rule mitigates catastrophic forgetting independently of activation saturation. Parameters: N=103 , K=10 , M=5 , L=5 , c=0.5 , α=50 . Results shown here are experiment-only.
Figure 6: Typical forgetting on the first task and transfer on the second task with the new initialization . In this setting, the SDGM procedure allows for no forgetting on the full range of task similarity, while allowing for the same transfer. The smaller transfer at big task similarity for LoRA + SDGM is due to long symmetric plateau, which size increases non-monotonically with c . Parameters : N= 103 , K= 10 , M= 5 , L= 5 , α=50 .
Figure 7: Additional results in the specialized regimes. a) Typical generalization error for full fine-tuning (red), LoRA (blue), and their SDGM-constrained variants. During Task 1 training, the symmetric plateau, corresponding to the absence of specialization, is located at logϵ∗≈−1.5 . b) Task 1 forgetting and c) Task 2 transfer as a function of teacher similarity c . d–e) Student overlaps with the teachers, as defined in Eqs.( 18 )–( 19 ), after the task switch for LoRA + SDGM (green) and standard + SDGM (yellow). The overlaps corresponding to frozen student directions remain constant throughout Task 2 training.
School of Computer Science and Engineering, Southeast University, Nanjing, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China · Huawei Technologies, Shenzhen, China
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China