LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models
Organizations: MMLA-org · Communication University of China (CUC), No. 1 Dingfuzhuang East Street, Chaoyang District, Beijing 100024, China
Abstract
Low-Rank Adaptive Residual Connections (LARC) give a frozen model a compact numerical state that can learn from feedback. The map adds a low-rank correction to a hidden representation. A slow state learns starting factors across tasks; a private fast state copies them, changes with feedback, and resets to the trained initialization. This report specifies an input-side realization of the numerical policy carrier in Memory-Mediated Learning Architecture and examines its factor-space dynamics and learning lifetime. We study a rank-4 input residual with 12,288 trainable parameters on a frozen MiniCPM5-1B-SFT substrate. In a four-candidate program-selection task, two feedback-gradient steps reduce expected query execution error by 24.65 and 36.65 percentage points relative to resetting to the respective trained static and post-adaptation initializations. These development results cover 16 parameter groups and three paired training seeds. A direct support-loss selection rule is much more accurate, reaching 0.78125% error. In a repository-balanced chronological replay of public continuous-integration jobs, retaining online updates raises half-Brier loss from 0.1274 to 0.1808. A fixed follow-up intervention records same-batch non-descent and inconsistent future benefit from shrinking updates. Together, the algebra and measurements distinguish residual capacity, adaptation relative to a starting point, and usefulness on later decisions.
Figures & tables
| State | What it contains | When it changes | What a fast reset does |
|---|---|---|---|
| Frozen substrate | Backbone and semantic adapter | Fixed during these experiments | Leaves it unchanged |
| Slow | Learned factors and version | One outer update after a batch closes | Uses it as the reference |
| Fast | Episode factors and private update state | Arrived-feedback update | Copies the bound |
| Explicit memory | Completed cases in the CI study | Real-label arrival under its own rule | Keeps its separate lifetime |
| Choice | Evaluated input LARC | Weight-space LoRA |
|---|---|---|
| Forward correction | at selected linear maps | |
| Linear overlap | Free factored | |
| Placement here | One hidden-to-hidden map before the full mixer | Depends on the chosen weight sites |
| Learning lifetime | Trained ; private, adapted ; reset to | Can use the same initialization, adaptation, and reset procedure |
| Objective | 2811 | 2812 | 2813 | Mean | Seed SD |
|---|---|---|---|---|---|
| Static | 21.90 | 31.47 | 34.53 | 29.30 | 6.59 |
| Adapted | 20.46 | 41.93 | 11.44 | 24.61 | 15.66 |
| Contrast | Mean | 99% group interval | 2811 | 2812 | 2813 |
|---|---|---|---|---|---|
| 4.69 | [1.19, 8.88] | 1.45 | -10.46 | 23.09 | |
| Static | 24.65 | [19.39, 29.26] | 31.43 | 20.16 | 22.35 |
| Static | 28.36 | [21.85, 34.93] | 37.76 | 26.78 | 20.55 |
| Adapted | 36.65 | [31.65, 41.20] | 40.52 | 18.64 | 50.78 |
| Adapted | 35.77 | [30.33, 41.01] | 40.21 | 19.97 | 47.13 |
| Objective | Real/keep | Real/reset | Sham/keep | Sham/reset | Text | Rule |
|---|---|---|---|---|---|---|
| Static | 29.30 | 53.95 | 57.67 | 53.95 | 61.78 | 0.78 |
| Adapted | 24.61 | 61.26 | 60.38 | 61.26 | 61.03 | 0.78 |
| Objective | Read state | Expected error | Greedy error | All query cases correct |
|---|---|---|---|---|
| Static | Reset | 53.95 | 47.76 | 42.71 |
| Static | Adapted state | 29.30 | 15.49 | 81.77 |
| Adapted | Reset | 61.26 | 61.72 | 30.21 |
| Adapted | Adapted state | 24.61 | 16.04 | 80.73 |
| Path | 2811 | 2812 | 2813 | Mean | Seed SD |
|---|---|---|---|---|---|
| P0M0 | 0.1291 | 0.1196 | 0.1372 | 0.1286 | 0.0088 |
| P1M0 | 0.1667 | 0.1542 | 0.1970 | 0.1726 | 0.0220 |
| P0M1 | 0.1162 | 0.1174 | 0.1487 | 0.1274 | 0.0184 |
| P1M1 | 0.1839 | 0.1658 | 0.1927 | 0.1808 | 0.0137 |
| PERM_M1 | 0.3193 | 0.2867 | 0.3854 | 0.3305 | 0.0503 |
| ORIGINAL | 0.1681 | 0.1646 | 0.1633 | 0.1653 | 0.0025 |
| Repository | Label | Jobs | Commits | Reset | Keep | Permuted | HEDGE4 |
|---|---|---|---|---|---|---|---|
| NumPy | All | 83 | 1 | 0.0300 | 0.0273 | 0.2656 | 0.0277 |
| NumPy | Success | 83 | 1 | 0.0300 | 0.0273 | 0.2656 | 0.0277 |
| pandas | All | 3238 | 71 | 0.2248 | 0.3343 | 0.3954 | 0.1772 |
| pandas | Success | 2310 | 63 | 0.0401 | 0.2048 | 0.3926 | 0.0702 |
| pandas | Failure | 32 | 12 | 0.9386 | 0.7644 | 0.3842 | 0.7128 |
| pandas | Cancelled | 704 | 18 | 0.8475 | 0.6987 | 0.4050 | 0.5847 |
| Aggregation | Reset | Keep | Permuted | HEDGE4 |
|---|---|---|---|---|
| Repository then commit (primary) | 0.1274 | 0.1808 | 0.3305 | 0.1024 |
| Commit (descriptive) | 0.2221 | 0.3300 | 0.3936 | 0.1751 |
| Job (descriptive) | 0.2329 | 0.3410 | 0.3934 | 0.1836 |
| Update path | 2811 | 2812 | 2813 | Mean | Seed SD |
|---|---|---|---|---|---|
| CARRY_1 | 0.3577 | 0.1925 | 0.3400 | 0.2967 | 0.0907 |
| CARRY_TENTH | 0.2847 | 0.3384 | 0.3382 | 0.3204 | 0.0310 |
| LATEST_1 | 0.2666 | 0.2488 | 0.3115 | 0.2756 | 0.0323 |
| LATEST_TENTH | 0.2223 | 0.1748 | 0.2713 | 0.2228 | 0.0482 |
| RESET | 0.1775 | 0.2214 | 0.2302 | 0.2097 | 0.0282 |
| HEDGE4 | 0.1679 | 0.1679 | 0.1679 | 0.1679 | 0.0000 |
| Seed | Before update | Original increment | Tenth increment | Original increment norm |
|---|---|---|---|---|
| 2811 | 0.513218 | 3.007320 | 0.102798 | 1.590172 |
| 2812 | 2.503274 | 2.838060 | 2.805870 | 7.709818 |
| 2813 | 0.721480 | 1.350831 | 0.139986 | 2.632058 |
| Study | Slow updates | Entry wall time (s) | Saved learning state |
|---|---|---|---|
| Episodic training and evaluation | 1,536 | 1,634.503 | Six slow initializations, version 256 |
| CI task fit and complete replay | 2,010 | 13,229.793 | Three slow initializations, version 926; six online states |
| CI prefix intervention | 0 | 1,058.808 | Six sessions containing all four active fast states |
| Operation | Calls | Positions per call | Ledger contribution |
|---|---|---|---|
| Mixer, including recomputation | 11,581 | 47,435,776 | |
| Selected-answer head | 6,195 | 1,585,920 | |
| Native reference | 1 | 4,096 | |
| Total | 49,025,792 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Index | Arithmetic recipe | List recipe |
|---|---|---|
| 0 | mul, add | filter_min, take |
| 1 | add, mul | take, filter_min |
| 2 | clip, mul, add | filter_min, sort, take, append |
| 3 | mul, clip, add | append, filter_min, sort, take |
| Artifact | Exact identity in accompanying manifest | Access at report preparation |
|---|---|---|
| MiniCPM5-1B-SFT | Revision a60b37f1 …; base file and tensor hashes | Public upstream revision |
| Frozen semantic | Original and exported file hashes; tensor identity | Local initialization package; not publicly hosted |
| Three static | Seed-specific original/export hashes | Local initialization package; not publicly hosted |
| Three adapted | Original checkpoint hashes and versions | Project archive; author access |
| Three CI | Original checkpoint hashes; parent lineage | Project archive; author access |
| Scientific implementation | Program-selection source revision 7a8e7b77 …; CI source revision f8a8ed95 … | Private project repository; author access |