cs.LGSep 29, 2026

Continual Learning of Dynamical Systems in Recurrent Neural Networks through Recyclable Unit Gating

Authors: Sima Hashemi, Daniel Durstewitz, Georgia Koppe

Organizations: Faculty of Mathematics and Computer Science, Interdisciplinary Center for Scientific Computing, Heidelberg University, Heidelberg, Germany · Department of Theoretical Neuroscience, Central Institute of Mental Health (CIMH), Medical Faculty Mannheim, Heidelberg University, Heidelberg, Germany · Hector Institute for AI in Psychiatry & Department of Psychiatry and Psychotherapy, CIMH, Medical Faculty Mannheim, Heidelberg University, Heidelberg, Germany · Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany

Abstract

Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continual DSR (cDSR) requires learning new systems while preserving previously learned dynamics, yet even small parameter updates in recurrent models can qualitatively alter their behavior over long autonomous rollouts. We benchmark established continual learning (CL) methods spanning parameter regularization, replay, and parameter isolation on the fully trainable and interpretable Almost-Linear RNN (AL-RNN). Parameter isolation preserves earlier dynamics most effectively, but excessive task-specific allocations can rapidly exhaust a fixed-size network. We therefore introduce Continually-Recyclable Unit-Gating (CRUG), which conserves capacity through compact allocation and forward transfer. Differentiable gates trained with an L0L_0-based penalty select task-specific units, while unused units are recycled for subsequent tasks. Directed connections allow later tasks to reuse earlier representations without affecting the dynamics of previously committed units. CRUG achieves the strongest reconstruction--capacity trade-off among the tested methods with zero forgetting and reliably learns a heterogeneous sequence of nonlinear and chaotic systems. Furthermore, we show that forward transfer is more pronounced and useful when tasks share similar underlying dynamics. Lastly, we demonstrate that CRUG's advantages extend beyond autonomous cDSR to sequential cognitive tasks.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 12, 2026cs.LG

Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction

Reconstructing nonlinear dynamical systems (DS) from data (DSR) is a fundamental challenge in science and engineering, but it inherently relies on sequential models. Recent breakthroughs for sequential models have produced algorithms that parallelize computation along sequence length TT, achieving logarithmic time complexity, O(log⁡T)\mathcal{O}(\log T). Since sequence lengths have been practically limited due to the linear runtime complexity O(T)\mathcal{O}(T) of classical backpropagation through time, this opens new avenues for DSR. This paper studies two prominent classes of parallel-in-time algorithms for this task, both of which leverage parallel associative scans as their core computational primitive. The first class comprises models with linear yet non-autonomous dynamics and a nonlinear readout, such as modern State Space Models (SSMs), while the second consists of general nonlinear models which can be parallelized using the DEER framework. We find that the linear training-time recurrence of the first class of models imposes limitations that often hinder learning of accurate nonlinear dynamics. To address this, we augment DEER with Generalized Teacher Forcing (GTF), a novel variant within the more general nonlinear framework that ensures stable and effective learning of nonlinear dynamics across arbitrary sequence lengths. Using GTF-DEER, we investigate the benefits of training on extremely long sequences (T>104T>10^4) for DSR. Our results show that access to such long trajectories significantly improves DSR if the data features long time scales. This work establishes GTF-DEER as a robust tool for data-driven discovery and underscores the largely untapped potential of long-sequence learning in modeling complex DS.
Jun 29, 2026cs.LG

Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management

We introduce Neural Subspace Reallocation (NSR), which reframes continual learning as memory management over parameter subspaces. Instead of treating Low-Rank Adaptation (LoRA) modules as disposable per-task adapters, NSR manages them as compressible, retrievable memory units on a frozen backbone through a recurring cycle: (1) compress learned LoRAs via SVD, (2) reserve them in a TaskKnowledgeBank, (3) recall related past LoRAs by embedding similarity to warm-start new or returning tasks, and (4) reallocate the active subspace accordingly, with distillation protecting prior tasks. We prove that in cyclic environments any memoryless allocation policy incurs cumulative regret Omega(T(M-1)Delta_switch) relative to a history-aware policy backed by the Bank (Theorem 1). Empirically, on Split-CIFAR-100 the Bank reduces cyclic recovery time by 10x, exactly as predicted, and on the heterogeneous 5-Datasets benchmark NSR achieves the highest accuracy and the least forgetting, about 9x closer to zero backward transfer than the memoryless heuristics. Crucially, we run a controlled study that isolates which component matters: holding the Bank fixed and varying only the allocation rule, we find that a simple similarity-based retrieval rule matches or beats a learned reinforcement-learning controller (recovering recurring tasks in 0 vs 1.8 steps and reaching equal accuracy). Our central, honest finding is therefore that the memory mechanism -- compression and similarity retrieval -- rather than a learned allocation policy, drives continual-learning performance under fixed capacity. A memory-budget analysis confirms the compressed Bank stays small -- 0.29 MB of parameter memory per task -- so a top-K retention cap bounds the total footprint while preserving fast recovery for retained tasks.
Sep 8, 2026cs.LG

Learning Length-Extrapolatable Recurrent Models

Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.