Context-dependent time-series prediction via HyperReservoirs
Authors: Kohei Tsuchiyama, Takatomo Mihana, Ryoichi Horisaki, André Röhm
Organizations: Department of Information Physics and Computing, Graduate School of Information Science and Technology, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan
Time series prediction is a common application of reservoir computing. When the training and testing time series data contains multiple dynamical regimes, because an underlying parameter is changing, or the data in fact consists of multiple distinct systems, simple application of the reservoir computing principle produces high prediction errors. Here, we propose a HyperReservoir as an extended model of reservoir computing especially designed for such cases. The HyperReservoir combines a main reservoir with a smaller context reservoir, where the latter modulates the output weights of the former. This structure resembles the hypernetworks from deep neural network literature. However, in contrast, HyperReservoirs retain the simple training via linear regression of standard reservoir computing. We compare the proposed architecture with a conventional ESN, in which context acts at the input, and a full-matrix Conceptor, in which context modulates the reservoir state space. We evaluate all three models on time-series prediction tasks based on Lorenz and Rössler systems, including for varying bifurcation parameters and time sampling scales. We find that the HyperReservoir achieves the lowest mean test error in all three tasks, and particularly outperforms conceptors on data that is sampled from the same attractor but at different time scales.
Figures & tables
Figure 1: Three mechanisms for introducing contextual information into a shared reservoir system. (a) In the context-input ESN, context enters through the external input while the recurrent operator and readout remain shared. (b) In the full-matrix Conceptor model, context selects a state-space operator Ck that filters the provisional reservoir state. (c) In the HyperReservoir, an additional context reservoir encodes the supplied context, and the bilinear term htH⊗htR allows this context representation to modulate the coefficients applied to the shared main-reservoir features. The recurrent matrices remain fixed in all three architectures.
Figure 2: Dynamical settings used for contextual future-state prediction. (a) Geometrically distinct Rössler and Lorenz attractors. (b) Two Rössler regimes with cR=3.5 and cR=5.7 . (c) The same Rössler attractor traversed at two temporal scales, νslow=0.5 and νfast=1.5 . Scaling the complete vector field preserves the continuous-time orbit geometry while changing the rate of traversal. Representative traces of the ω component illustrate the resulting temporal difference.
Figure 3: Future-state prediction for the three dynamical settings introduced in Sec. III . (a) Geometrically distinct Lorenz and Rössler systems. (b) Two Rössler regimes with cR=3.5 and cR=5.7 . (c) The same Rössler attractor traversed at different temporal scales, νslow=0.5 and νfast=1.5 . Test component-normalized mean squared error (cNMSE) is shown for the context-input ESN, full-matrix Conceptor, and augmented HyperReservoir. Individual markers denote reservoir realizations; horizontal markers and error bars show the mean and sample standard deviation. The vertical axes are logarithmic, and lower cNMSE indicates more accurate prediction.
Figure 4: Analysis of contextual prediction on a shared attractor. (a,b) One-step prediction errors, eχ,t=χ^t+1−χt+1 , for the slow ν=0.5 and fast ν=1.5 regimes. The same representative reservoir realization and evaluation window are used for all models; the vertical scales differ between panels for visibility. (c) Dependence on the supplied context, quantified by the ratio Rctx between correct- and wrong-context prediction errors. Lower values indicate a larger increase in error after context replacement. (d) Frobenius cosine similarity between the slow and fast Conceptors for the three reservoir realizations. Values close to unity indicate strong alignment of the two operators.
Figure 5: Structure and dimensionality of the HyperReservoir readout. (a) Three readout constructions. The Concat model uses the main- and context-reservoir states additively, the Strict model uses only their bilinear interaction, and the Augmented model combines both. (b) Test cNMSE for the three readout constructions across the dynamical settings introduced in Sec. III . Individual markers denote reservoir realizations, and error bars show the mean and sample standard deviation. The Lorenz–Rössler setting uses N=110,M=10 , whereas the related-Rössler and same-attractor settings use N=115,M=5 . (c) Test cNMSE of the Augmented HyperReservoir for the related-Rössler setting as the context-reservoir dimension M is varied while N+M=120 . The dashed line denotes the corresponding context-input ESN result. (d) Number of fitted output coefficients as M is varied under the same constraint.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Quantity
Value
Data
data seed
0
model seeds
{0,1,2}
training trajectories
256
validation trajectories
64
test trajectories
64
samples per trajectory
1000
Appendix
Table S1: Common numerical and model settings used in the three contextual prediction tasks. The model seed changes the random recurrent realization while the generated datasets remain fixed.
Model
Main reservoir
Context reservoir / selector
Context-dependent operation
Context-input ESN
observation + context
none
input forcing
full-matrix Conceptor
observation only
context selects Ck
state-space filtering
HyperReservoir
observation only
context only
readout modulation
Appendix
Table S2: Architecture-specific routing of the observed dynamical state and context.
Task
Model
Seed 0
Seed 1
Seed 2
Mean ± sample SD
Distinct
ESN
4.5365×10−2
4.4232×10−2
4.4425×10−2
(4.4674±0.0606)×10−2
Conceptor
4.7693×10−2
4.2918×10−2
4.1809×10−2
(4.4140±0.3126)×10−2
HyperReservoir
3.2950×10−2
3.9176×10−2
3.1146×10−2
(3.4424±0.4213)×10−2
Related
ESN
4.6593×10−4
5.7953×10−4
3.4698×10−4
(4.6415±1.1629)×10−4
Conceptor
4.1522×10−3
3.5667×10−3
8.5255×10−3
(5.4148±2.7098)×10−3
HyperReservoir
8.7665×10−5
7.3333×10−5
7.5406×10−5
(7.8801±0.7746)×10−5
Appendix
Table S3: Seedwise test cNMSE. The final column gives the arithmetic mean and sample standard deviation over the three random reservoir realizations.
M
Distinct
Related
Same attractor
2
(3.9722±0.3106)×10−2
(1.8801±0.3013)×10−4
(1.6162±0.8627)×10−3
5
(3.5676±0.1228)×10−2
(7.8801±0.7746)×10−5
(6.4416±1.0439)×10−4
10
(3.4424±0.4213)×10−2
(1.1040±0.1318)×10−4
(6.6432±1.4084)×10−4
20
(3.5997±0.2987)×10−2
(1.4243±0.1530)×10−4
(8.4349±3.8616)×10−4
40
(3.7962±0.3321)×10−2
(2.5571±1.2052)×10−4
(7.7254±1.5993)×10−4
Appendix
Table S4: Test cNMSE for the augmented HyperReservoir as the context-reservoir dimension M is varied under the fixed recurrent-state budget N+M=120 . Values are the mean and sample standard deviation across three reservoir seeds.
Readout
Distinct
Related
Same attractor
Concat
(4.4320±0.3341)×10−2
(2.0709±0.3444)×10−4
(2.0795±0.1217)×10−3
Strict multiplicative
(5.3592±0.6946)×10−2
(9.1011±2.2018)×10−2
(1.9783±0.9186)×10−3
Augmented
(3.4424±0.4213)×10−2
(7.8801±0.7746)×10−5
(6.4416±1.0439)×10−4
Appendix
Table S5: Readout comparison using the validation-selected N and M of the augmented HyperReservoir for each task. Values are mean test cNMSE and sample standard deviation across three random reservoir realizations.
Figure S1: Counterfactual context dependence in the same-attractor temporal-scale task. Mean cNMSE matrices for (a) the context-input ESN, (b) the full-matrix Conceptor, and (c) the HyperReservoir. Rows indicate the true temporal-scale context and columns indicate the supplied context, with ν∈{0.5,1.5} . Diagonal entries therefore correspond to the correct context, whereas off-diagonal entries correspond to counterfactual context replacement. Each matrix is averaged over the three reservoir realizations after task-level validation-based model selection. A common color scale is used across the three architectures.
Allocation
Model
Feature dimension
Fitted readout coefficients
N=110,M=10
Concat
120
363
Strict
1100
3303
Augmented
1220
3663
N=115,M=5
Concat
120
363
Strict
575
1728
Augmented
695
2088
Appendix
Table S6: Readout feature dimensions and fitted output coefficients for the recurrent-state allocations used in the readout ablation. The Conceptor storage entry is not included because it is separate from the shared linear readout.
M
N
Faug
Paug
2
118
356
1071
5
115
695
2088
10
110
1220
3663
20
100
2120
6363
40
80
3320
9963
Appendix
Table S7: HyperReservoir feature dimension and fitted readout size across the context-dimension sweep under N+M=120 .
Reservoir computing has emerged as an efficient machine learning framework for predicting time series generated by dynamical systems. In contrast to other machine and deep learning approaches, a reservoir computing trains only the output layer via linear regression, leaving the reservoir (recurrent layer) untrained. This simplification makes reservoir computers easier to train and more amenable to experimentation. However, because current reservoirs consist of networks of randomly connected nodes and require the optimization of numerous hyperparameters, a framework that precisely explains how reservoir computing operates and how it can be optimized remains missing. Here, we propose a frequency-based reservoir inspired by the brain's oscillatory dynamics and its hierarchy of timescales. The frequency-based reservoir can be interpreted as an ensemble of independent oscillatory units, each processing a portion of the input's frequency content. This allows us to understand the reservoir's internal behavior by modeling it as a single unit driven by an external input. Borrowing from the theory of a nonlinear oscillator forced by complex periodic inputs, we found that units of the frequency-based reservoir selectively amplify and store specific input frequencies, which are then used for prediction. The frequency-based reservoir performs as well as or better than equivalent random reservoirs. Furthermore, the frequency-based approach can be optimized to improve short-term prediction, a property that random reservoirs lack. Finally, we show that the frequency-based reservoir can also predict complex spatiotemporal dynamics. Our results show that reservoir computing can be designed using brain properties and theoretical insights borrowed from the physics of forced nonlinear oscillators.
Arthur S Powanwe
Department of Mathematics, Western University, London, ON, Canada
We present an adaptive reservoir computing framework for the CTF-4-Science Lorenz benchmark, which evaluates machine learning models across twelve distinct tasks spanning five qualitatively different scenarios: baseline forecasting, noisy signal reconstruction, forecasting under noise, few-shot learning, and parametric generalization. Rather than applying a uniform inference strategy, we tailor the training and prediction procedure of Echo State Networks (ESNs) to the specific demands of each evaluation scenario. Our key contributions are fourfold: (1) exact reservoir state synchronization that eliminates warmup approximation error in short-time prediction; (2) histogram-guided candidate selection that directly optimizes the long-time ergodic evaluation metric; (3) multi-seed reservoir search for few-shot regimes with severely limited training data; and (4) sequential multi-sequence training that resolves state-distribution mismatch in parametric generalization tasks. The proposed framework achieves a score of 74.91 on the public benchmark leaderboard, demonstrating that carefully adapted reservoir computing constitutes a competitive and computationally efficient approach for diverse chaotic system modeling challenges.
Reservoir computing typically relies on large, randomly generated reservoirs, enabling simple, often linear readouts. Over the past two decades, most constructions have exploited the freedom to select the reservoir, constrained primarily by stability conditions based on state contraction or memory capacity. However, these designs are largely independent of the input data and learning objective, resulting in a trial-and-error methodology driven by randomness. In high dimensions, the reservoir acts as a random embedding of the input history, implicitly relying on Johnson--Lindenstrauss--type concentration phenomena to preserve information. In contrast, we develop reservoir design principles from a geometric perspective for inputs generated by deterministic dynamical systems. Rather than relying on random embeddings, we require reservoir state increments to align within a cone around an input-determined vector subspace, and prove that such a cone concentration reduces ridge-regression training error. When the cone angle is small, the variance of reservoir states concentrates in the input-determined subspace, improving conditioning of the empirical second-moment matrix and strengthening alignment between dominant covariance directions and the state-target cross-covariance. For echo state networks, we provide a constructive approach to reservoir design. The reservoir matrix is chosen so that associated Krylov-chain directions remain nearly closed within an input-determined subspace while permitting controlled mixing in its orthogonal complement. We also provide a spectral diagnostic for ridge regression training that identifies when reservoir geometry concentrates predictive information into a few dominant covariance modes and when ``spectral pollution'' inhibits forecasting. Numerical experiments demonstrate consistent performance gains over arbitrary reservoir constructions.
G Manjunath, Juan-Pablo Ortega, Alma van der Merwe
University of Pretoria. Department of Mathematics and Applied Mathematics. Pretoria 0002, South Africa. · Nanyang Technological University. Division of Mathematical Sciences. School of Physical and Mathematical Sciences. Singapore.