From Redundancy to Minimality: Fixed-Point-Guided Hierarchical Reduction of Learned Piecewise-Linear Dynamics
Authors: Hiroto Tamura, Gouhei Tanaka
Organizations: Graduate School of Informatics, Osaka Metropolitan University, Osaka, Japan · International Research Center for Neurointelligence (IRCN), The University of Tokyo, Tokyo, Japan · Graduate School of Engineering, Nagoya Institute of Technology, Nagoya, Japan
Understanding a nonlinear dynamical system from time series requires not only reproducing its trajectories, but also identifying a simple representation that preserves its essential dynamical structure. Almost-linear recurrent neural networks (AL-RNNs) are piecewise-linear RNNs in which only a subset of units use ReLU nonlinearities, so that nonlinear capacity is explicitly controlled by the number of ReLU units. Their activation patterns define linear regions, represented as symbols, whose observed transitions form a symbolic transition graph. However, directly training AL-RNNs with few ReLU units to realize minimal dynamical representations can be unreliable. We ask whether an AL-RNN with more ReLU units can instead be trained first and systematically reduced to a minimal dynamical representation. We introduce a fixed-point-guided hierarchical reduction procedure that progressively linearizes selected ReLU units, merging neighboring linear regions and graph nodes while preserving distinct symbols containing fixed points (FPs). The resulting reduction tree defines a hierarchy of progressively simpler candidates. Each reduced candidate is initialized from the parent parameters and retrained under guidance from the parent dynamics. We also prove that reproducing Q distinct fixed points requires at least Q FP-containing symbols, providing a certificate of symbol-level minimality when this bound is attained. On the 3-scroll Chua system, direct training with the theoretical minimum of three ReLU units achieves high-fidelity minimal realizations in only 20% of seeds, whereas our learn-reduce-retrain strategy increases the seed-macro success rate to approximately 71% at the same final nonlinear capacity. These results show that redundant nonlinear capacity can serve as a scaffold for discovering and realizing minimal dynamical representations.
Figures & tables
Figure 1: Proposed fixed-point-guided hierarchical reduction. (a) FP and FP-free symbols in the piecewise-linear state-space partition. (b) Symbolic transition graph constructed from visited ReLU activation patterns. (c) Region merging induced by ReLU linearization. (d) Reduction-tree construction under the fixed-point-preservation constraint.
Figure 2: Main results for the 3-scroll Chua system. (a) Direct training becomes more reliable as the number of ReLU units P increases. (b) Hierarchical reduction reaches ∣ΣD∣=5 and PD=3 , attaining both the symbol-level and nonlinear-capacity lower bounds. (c) Learning large first provides a more reliable route to minimal models than direct training at the minimal capacity: seed-macro success rates are substantially higher after reduction and retraining (left), while the resulting PD=3 models also exhibit substantially lower Estsp than directly trained P=3 models (right). The P=10 parents serve as a reference for the geometric fidelity attainable before reduction.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Pparent
2Pparent
Admissible nodes
Tree time [s]
10
1,024
476 [382–586]
0.2 [0.1–0.3]
11
2,048
1,133 [861–1,270]
0.5 [0.4–0.7]
12
4,096
2,324 [1,936–2,685]
1.4 [1.2–1.7]
13
8,192
5,145 [3,554–5,498]
3.3 [2.7–4.3]
14
16,384
10,616 [8,543–11,779]
10.2 [8.1–11.5]
Appendix
Table A1: Empirical scaling of exhaustive reduction-tree construction. Values are medians over 30 parent models, with interquartile ranges in brackets.
Figure A1: Ablation of reduction-guided retraining on the 3-scroll Chua system. All methods use the same selected five-symbol reduction candidates and parent-derived initialization. Top: parent-balanced seed-macro success rates for fixed-point recovery, minimal realization, and high-fidelity minimal realization ( Sfid=1 , requiring Estsp<2.0 ). Full-preactivation guidance achieves 77% , 74% , and 73% , respectively, compared with 19% , 14% , and 14% without parent-trajectory guidance. Symbol and retained-preactivation guidance also substantially improve success rates. Bottom left: post-transient state-space error Estsp ; the dashed line marks the Estsp=2.0 high-fidelity threshold. Bottom right: post-transient spectral Hellinger error EH , shown as a complementary dynamical diagnostic. Points denote individual candidates with valid finite measurements, and boxplots summarize the corresponding parent-balanced distributions.
Figure A2: Representative autonomous 3-scroll Chua trajectories across state-space reconstruction errors. Gray trajectories show the ground-truth attractor, and blue trajectories show the corresponding autonomous model outputs. From top to bottom, the rows contain four representative models with Estsp≈1.5 , 2.0 , 3.0 , and 4.0 , respectively; the exact Estsp value of each model is shown above its panel. Models with Estsp<2.0 generally preserve the overall 3-scroll geometry, whereas larger errors increasingly include distorted, imbalanced, or incomplete attractors. These examples illustrate the practical motivation for using Estsp<2.0 as the high-fidelity criterion.
Figure A3: Sensitivity of high-fidelity minimal-realization rates to the state-space-error threshold on the Chua system. The curves show H(τ)=Pr[Smin=1∧Estsp<τ] for direct minimal-capacity training, retraining without parent-trajectory guidance, and the three guided retraining variants. The vertical dashed line marks the threshold Estsp=2.0 used in the main analysis. Reduction-guided retraining maintains substantially higher success rates than direct P=3 training over a broad range of thresholds, and the guided curves are already close to their plateaus by τ=2.0 .
Figure A4: Reorganization of non-chain quotient candidates into a common five-node chain after retraining. The left column shows representative selected quotient graphs before retraining, including a two-cycle topology without a leaf and one-cycle-plus-tail topologies. The right column shows the corresponding autonomous transition graphs after reduction-guided retraining. Despite the different predicted quotient structures, all three retrained models realize Qvis=5 and ∣Σ∣=5 , with low post-transient state-space error ( Estsp=1.17–1.28 ), and converge to the same path-like five-node organization. Stars indicate symbols associated with admissible fixed points.
We investigate a structure-first approach to dynamical learning in which the organization of stateful interactions is prescribed explicitly rather than left entirely to a generic recurrent parameterization. We introduce causal recurrent units built from an ordered sequence of local, state-modulated transformations. The construction is motivated by wave-based interaction models, but the units studied here do not impose scattering, passivity, or energy-balance constraints. Because fixed recurrent dynamics, designed reservoir topologies, readout-only learning, and recurrent depth are already well established, the empirical question is deliberately narrower: does the proposed interaction organization provide a useful inductive bias under controlled computational conditions? We compare a one-layer structured model, a two-layer structured model, and a generic echo-state network (ESN), all with 12 recurrent states and the same strictly linear ridge readout. Each model family receives the same random-search budget on calibration data that are disjoint from the final test data, after which the selected hyperparameters are frozen. On a custom nonlinear identification task, the one-layer structured model attains a mean validation NMSE of 2.76 x 10^{-4}, compared with 3.19 x 10^{-4} for the two-layer model and 3.94 x 10^{-4} for the ESN. On NARMA10 the ordering reverses: the ESN attains 0.312, compared with 0.348 and 0.357 for the one- and two-layer structured models. Thus, the proposed organization can be competitive and advantageous on one task, but it is not universally superior; moreover, recurrent depth does not provide a systematic benefit under matched state dimension. The results support a task-dependent interpretation of structural inductive bias and position the present architecture as a controlled precursor to stronger wave- and system-theoretic constructions.
Augusto Sarti
Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB), Politecnico di Milano, Piazza L. Da Vinci 32, 20133 Italy.
Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continual DSR (cDSR) requires learning new systems while preserving previously learned dynamics, yet even small parameter updates in recurrent models can qualitatively alter their behavior over long autonomous rollouts. We benchmark established continual learning (CL) methods spanning parameter regularization, replay, and parameter isolation on the fully trainable and interpretable Almost-Linear RNN (AL-RNN). Parameter isolation preserves earlier dynamics most effectively, but excessive task-specific allocations can rapidly exhaust a fixed-size network. We therefore introduce Continually-Recyclable Unit-Gating (CRUG), which conserves capacity through compact allocation and forward transfer. Differentiable gates trained with an L0-based penalty select task-specific units, while unused units are recycled for subsequent tasks. Directed connections allow later tasks to reuse earlier representations without affecting the dynamics of previously committed units. CRUG achieves the strongest reconstruction--capacity trade-off among the tested methods with zero forgetting and reliably learns a heterogeneous sequence of nonlinear and chaotic systems. Furthermore, we show that forward transfer is more pronounced and useful when tasks share similar underlying dynamics. Lastly, we demonstrate that CRUG's advantages extend beyond autonomous cDSR to sequential cognitive tasks.
Sima Hashemi, Daniel Durstewitz, Georgia Koppe
Faculty of Mathematics and Computer Science, Interdisciplinary Center for Scientific Computing, Heidelberg University, Heidelberg, Germany · Department of Theoretical Neuroscience, Central Institute of Mental Health (CIMH), Medical Faculty Mannheim, Heidelberg University, Heidelberg, Germany · Hector Institute for AI in Psychiatry & Department of Psychiatry and Psychotherapy, CIMH, Medical Faculty Mannheim, Heidelberg University, Heidelberg, Germany +1
Learning in recurrent neural networks can fundamentally reshape their underlying dynamics, transforming initially chaotic activity into stable task-dependent behavior. We develop a non-equilibrium dynamical mean-field theory(DMFT) to describe this transition during learning. We show that a slow feedback-driven learning process generates an evolving effective feedback strength that drives the network through a transition from chaotic to stable dynamics defined by a bifurcation of the DMFT solution. By deriving the two-time correlation function throughout learning, we identify a critical feedback strength and a corresponding learning rate dependent critical time separating these regimes. The transition arises from the progressive deformation of an effective dynamical landscape by the growing learned feedback structure. Starting from the untrained state, the theory predicts the time evolution of the network output during training and shows quantitative agreement with numerical simulations.
Varun Vaidya
Department of Physics, University of South Dakota, Vermillion 57069, USA