Low-Rank Adaptive Residual Connections (LARC) give a frozen model a compact numerical state that can learn from feedback. The map h+BAh adds a low-rank correction to a hidden representation. A slow state ρ learns starting factors across tasks; a private fast state Φ copies them, changes with feedback, and resets to the trained initialization. This report specifies an input-side realization of the numerical policy carrier in Memory-Mediated Learning Architecture and examines its factor-space dynamics and learning lifetime. We study a rank-4 input residual with 12,288 trainable parameters on a frozen MiniCPM5-1B-SFT substrate. In a four-candidate program-selection task, two feedback-gradient steps reduce expected query execution error by 24.65 and 36.65 percentage points relative to resetting to the respective trained static and post-adaptation initializations. These development results cover 16 parameter groups and three paired training seeds. A direct support-loss selection rule is much more accurate, reaching 0.78125% error. In a repository-balanced chronological replay of public continuous-integration jobs, retaining online updates raises half-Brier loss from 0.1274 to 0.1808. A fixed follow-up intervention records same-batch non-descent and inconsistent future benefit from shrinking updates. Together, the algebra and measurements distinguish residual capacity, adaptation relative to a starting point, and usefulness on later decisions.
Figures & tables
Figure 2: The input residual and its learning path. The two factors read a four-dimensional projection and write a correction back into the full hidden representation. The orange path denotes backpropagation to the factors through the frozen downstream computation. Zero B makes the initial forward map the identity; an episode reset later restores the trained factors.
State
What it contains
When it changes
What a fast reset does
Frozen substrate
Backbone and semantic adapter ω
Fixed during these experiments
Leaves it unchanged
Slow ρv
Learned factors (Aρ,Bρ) and version
One outer update after a batch closes
Uses it as the reference
Fast Φe,k
Episode factors and private update state
Arrived-feedback update
Copies the bound ρv
Explicit memory M
Completed cases in the CI study
Real-label arrival under its own rule
Keeps its separate lifetime
Table 1: The states read by the evaluated system. Explicit memory is disabled in the episodic study. In CI it provides the same causal case history to the compared policy states; resetting the numerical residual does not erase that history.
Figure 3: Slow learning and fast adaptation. Every episode in a batch begins at the same slow version and follows its own support trajectory. The objective chooses whether its query is evaluated at ρv or at the adapted factors. Query gradients are then aggregated into one slow update. At use time, the trained slow state stays fixed; fast resets and checkpoint restoration perform different operations.
Choice
Evaluated input LARC
Weight-space LoRA
Forward correction
Fθ((I+BA)h)
(W+UV)h at selected linear maps
Linear overlap
ΔW=(WB)A
Free factored ΔW=UV
Placement here
One hidden-to-hidden map before the full mixer
Depends on the chosen weight sites
Learning lifetime
Trained ρ ; private, adapted Φ ; reset to ρ
Can use the same initialization, adaptation, and reset procedure
Table 2: Placement and lifetime are separate choices. The linear overlap is given by Equation ( 6 ); the evaluated LARC precedes the full nonlinear mixer.
Figure 4: Feedback binding crossed with state retention. Both feedback branches execute two updates. The reset read then restores its own trained initialization immediately before the query. With the remaining read inputs fixed, G measures the effect of retaining real-feedback state, while D compares this effect with independently permuted feedback.
Objective
2811
2812
2813
Mean
Seed SD
Static
21.90
31.47
34.53
29.30
6.59
Adapted
20.46
41.93
11.44
24.61
15.66
Table 3: Post-adaptation development error (%), lower is better. Each column evaluates the same 16 parameter groups. Seed SD is the sample standard deviation across the three trained initializations.
Contrast
Mean
99% group interval
2811
2812
2813
Δobj
4.69
[1.19, 8.88]
1.45
-10.46
23.09
Static G
24.65
[19.39, 29.26]
31.43
20.16
22.35
Static D
28.36
[21.85, 34.93]
37.76
26.78
20.55
Adapted G
36.65
[31.65, 41.20]
40.52
18.64
50.78
Adapted D
35.77
[30.33, 41.01]
40.21
19.97
47.13
Table 4: Paired gains in percentage points, positive is favorable. The bootstrap unit is a complete parameter group. Training seeds are kept paired within each resample.
Objective
Real/keep
Real/reset
Sham/keep
Sham/reset
Text
Rule
Static
29.30
53.95
57.67
53.95
61.78
0.78
Adapted
24.61
61.26
60.38
61.26
61.03
0.78
Table 5: Descriptive development errors (%). Text renders the program-to-feedback mapping in the prompt at fixed ρ . Rule chooses the lowest-support-error program, with first-index tie breaking. All methods receive the same external support information.
Figure 5: Two questions in the episodic study. Left: each line joins one paired initialization under static and adapted training; the dashed black line joins their means. Right: the five recorded gains and 99% task-group intervals. The adapted objective improves the mean query error, while both objectives support useful real-feedback adaptation. The intervals retain the three trained seeds in every resample.
Objective
Read state
Expected error
Greedy error
All query cases correct
Static
Reset
53.95
47.76
42.71
Static
Adapted state
29.30
15.49
81.77
Adapted
Reset
61.26
61.72
30.21
Adapted
Adapted state
24.61
16.04
80.73
Table 6: Additional descriptive development measures (%), reconstructed from saved predictions. Greedy error evaluates the highest-probability supplied program; the last column is the fraction of episodes where that program passes all 20 query inputs. Group/family weights and paired seeds match the primary analysis. These are secondary measures, not free-form generation.
Path
2811
2812
2813
Mean
Seed SD
P0M0
0.1291
0.1196
0.1372
0.1286
0.0088
P1M0
0.1667
0.1542
0.1970
0.1726
0.0220
P0M1
0.1162
0.1174
0.1487
0.1274
0.0184
P1M1
0.1839
0.1658
0.1927
0.1808
0.0137
PERM_M1
0.3193
0.2867
0.3854
0.3305
0.0503
ORIGINAL
0.1681
0.1646
0.1633
0.1653
0.0025
Table 7: Full CI replay, grouped half-Brier loss. PERM_M1 uses permuted fast-update labels and real memory. ORIGINAL retains the pre-CI slow initialization; CALIBRATED adds a four-class bias fitted on historical data. Both controls also receive historical and online fast updates.
Repository
Label
Jobs
Commits
Reset
Keep
Permuted
HEDGE4
NumPy
All
83
1
0.0300
0.0273
0.2656
0.0277
NumPy
Success
83
1
0.0300
0.0273
0.2656
0.0277
pandas
All
3238
71
0.2248
0.3343
0.3954
0.1772
pandas
Success
2310
63
0.0401
0.2048
0.3926
0.0702
pandas
Failure
32
12
0.9386
0.7644
0.3842
0.7128
pandas
Cancelled
704
18
0.8475
0.6987
0.4050
0.5847
Table 8: Saved CI half-Brier losses by repository and observed label, averaged over the three seeds. Rows labeled All weight commits equally within a repository; individual-label rows weight jobs equally within that label. Reset, Keep, and Permuted all use real memory. Label strata are descriptive and overlap in their commit membership.
Aggregation
Reset
Keep
Permuted
HEDGE4
Repository then commit (primary)
0.1274
0.1808
0.3305
0.1024
Commit (descriptive)
0.2221
0.3300
0.3936
0.1751
Job (descriptive)
0.2329
0.3410
0.3934
0.1836
Table 9: Weighting sensitivity on the same saved 3,321 predictions per seed. The first row remains the registered primary measure. The other rows are descriptive; neither introduces new jobs, retraining, or a replacement success rule.
Update path
2811
2812
2813
Mean
Seed SD
CARRY_1
0.3577
0.1925
0.3400
0.2967
0.0907
CARRY_TENTH
0.2847
0.3384
0.3382
0.3204
0.0310
LATEST_1
0.2666
0.2488
0.3115
0.2756
0.0323
LATEST_TENTH
0.2223
0.1748
0.2713
0.2228
0.0482
RESET
0.1775
0.2214
0.2302
0.2097
0.0282
HEDGE4
0.1679
0.1679
0.1679
0.1679
0.0000
Table 10: Fixed-prefix diagnostic, half-Brier loss. The primary comparison is CARRY_1 minus CARRY_TENTH . All paths use the same prefix. Their losses use a different evaluation population from Table 7 .
Seed
Before update
Original increment
Tenth increment
Original increment norm
2811
0.513218
3.007320
0.102798
1.590172
2812
2.503274
2.838060
2.805870
7.709818
2813
0.721480
1.350831
0.139986
2.632058
Table 11: Cross-entropy on the same first pandas feedback batch, before and after the numerical update. Increment norm is the Euclidean norm across the two factor tensors.
Figure 6: Three views of online learning. (a) The full 3,321-job replay compares real-memory paths at reset, retained real-feedback state, permuted-feedback state, and HEDGE4. (b) Same-batch CE change on the first eight pandas feedback jobs: positive values indicate non-descent. (c) The fixed 736-job diagnostic prefix: original carried-update loss minus the smaller carried-update loss, separately for each seed. The latter two panels describe a diagnostic on the observed queue; their populations and losses are distinct from panel (a).
Study
Slow updates
Entry wall time (s)
Saved learning state
Episodic training and evaluation
1,536
1,634.503
Six slow initializations, version 256
CI task fit and complete replay
2,010
13,229.793
Three slow initializations, version 926; six online states
CI prefix intervention
0
1,058.808
Six sessions containing all four active fast states
Table 12: Recorded execution costs. Wall times include model loading, evaluation, CPU work inside the entry, saving, and cleanup. The CI total includes two stopped preflights, an interrupted partial replay, and its completed evaluation. They are not pure GPU kernel times.
Operation
Calls
Positions per call
Ledger contribution
Mixer, including recomputation
11,581
8×512
47,435,776
Selected-answer head
6,195
8×32
1,585,920
Native reference
1
8×512
4,096
Total
49,025,792
Table 13: Decomposition of the episodic physical-position ledger. Head and native-reference charges explain the 1,590,016 positions beyond the mixer subtotal. Backward calls are counted separately.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Index
Arithmetic recipe
List recipe
0
mul, add
filter_min, take
1
add, mul
take, filter_min
2
clip, mul, add
filter_min, sort, take, append
3
mul, clip, add
append, filter_min, sort, take
Appendix
Table 14: All eight fixed recipes. A family’s four recipes instantiate the four candidate programs using its public parameter values.
Artifact
Exact identity in accompanying manifest
Access at report preparation
MiniCPM5-1B-SFT
Revision a60b37f1 …; base file and tensor hashes
Public upstream revision
Frozen semantic ω
Original and exported file hashes; tensor identity
Local initialization package; not publicly hosted
Three static ρ256
Seed-specific original/export hashes
Local initialization package; not publicly hosted
Three adapted ρ256
Original checkpoint hashes and versions
Project archive; author access
Three CI ρ926
Original checkpoint hashes; parent lineage
Project archive; author access
Scientific implementation
Program-selection source revision 7a8e7b77 …; CI source revision f8a8ed95 …
Private project repository; author access
Appendix
Table 15: Artifact availability and identity. An identity record does not imply that a private artifact is publicly hosted. The frozen upstream model retains its own distribution terms.
Reservoir computing (RC) trains only a linear readout over a fixed recurrent layer, making it fast and data-efficient for online prediction. However, a static reservoir degrades under system drift, readout-only adaptation is then insufficient, and unconstrained reservoir adaptation can destroy the echo-state and incremental stability properties that make RC reliable. This paper proposes LoRA-RC, which adapts the recurrent matrix through a low-rank correction driven by streaming prediction errors. The base reservoir and adaptation bases are fixed offline; a small core matrix is adapted online, projected onto a spectral-norm ball, and low-pass filtered at each step. The projection guarantees that every applied recurrent matrix remains within a certified contraction set, and an incremental input-to-state stability bound is established for the reservoir along each online adaptation path, with path-independent rate and gain. On a Lorenz system with an abrupt parameter drift, LoRA-RC cuts post-drift prediction error by 56% versus a fixed RC and 51% versus readout-only adaptation; ablations over 20 seeds show that removing the projection inflates this error by more than a factor of 40.
Wenbin Wan
Department of Mechanical Engineering, University of New Mexico, Albuquerque, NM 87131, USA
Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.
Department of Mathematics, Alma Mater Studiorum – Università di Bologna, Piazza di Porta San Donato 5, 40126 Bologna, Italy · Gatsby Computational Neuroscience Unit, University College London · Donders Centre for Neuroscience, Radboud University, Nijmegen, The Netherlands
Low-Rank Adaptation (LoRA) methods have emerged as crucial techniques for adapting large pre-trained models to downstream tasks under computational and memory constraints. However, they face a fundamental challenge in balancing task-specific performance gains against catastrophic forgetting of pre-trained knowledge, where existing methods provide inconsistent recommendations. This paper presents a comprehensive analysis of the performance-forgetting trade-offs inherent in low-rank adaptation using principal components of weight matrices as initialization. Our investigation reveals that fine-tuning intermediate components leads to better balance and robustness to high learning rates than first (PiSSA) and last (MiLoRA) components in existing work. Building on these findings, we provide practical guidelines for initialization of LoRA methods to balance the performance-forgetting trade-off. In a thorough empirical study on a variety of computer vision and NLP tasks we confirm that these guidelines achieve high accuracy and reduced forgetting.
Alessio Quercia, Arya Bangun, Ira Assent +1
IAS-8, Forschungszentrum Juelich, Juelich, Germany · Department of Computer Science, RWTH Aachen University, Aachen, Germany · Aarhus University, Aarhus, Denmark