Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-k nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model (k=0) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Increasing nonlinear innovation rank providing no measured benefits. This result is obtained on a fresh test split, which will be only opened after all models and analyses are frozen. The linear model requires fewer recurrent parameters and less computation than GRU, but performs comparably. These results suggest that, for this memory dependent control task, nonlinear representation learning around a simple linear memory mechanism can be sufficient, and that nonlinear recurrent dynamics are not necessarily required. These conclusions are limited to the simulated task, teacher, state dimension, and blackout horizon considered here, and to policies whose observation encoder and action head remain nonlinear.
Figures & tables
Figure 1: Why the task needs memory. (a) The two shots in a pair reach the same point when tracking stops, then continue towards opposite sides of the goal. The defender must remember which way the puck was moving. (b) Tracking stops after 100 ms, but the policy keeps acting. We test blackouts from 0 to 500 ms; 400 ms is the main comparison. (c) At 400 ms, the teacher clearly outperforms the policy without memory. The teacher also performs much worse if its state is erased when tracking stops. Both comparisons use the same 225 shots. Ranges show paired-bootstrap 95% confidence intervals. Erasing state has no effect when tracking remains available.
Figure 2: Structured recurrent student family. The public observation is encoded once and supplied both to the recurrent update and the nonlinear action head. The solid blue branch is the shared stable diagonal linear backbone. The dashed orange branch is the optional rank- k nonlinear innovation; it is absent for k=0 . All variants carry the same 64-value state and previous two-dimensional action. Only the state-dependent recurrent correction is rank controlled: the observation encoder and action head remain nonlinear in every variant.
Figure 3: Principal behavioural results. (a) The feed-forward student deteriorates sharply when the puck is hidden. A ten-step observation stack helps, while the structured k=0 student remains close to the teacher through the main 400 ms range. (b) On an expanded 90–100% scale, all four structured students and GRU-64 remain close to the teacher. The 500 ms point is an extrapolation beyond the main range. Student curves are means over five matched seeds; bars show paired-bootstrap 95% confidence intervals. The teacher is one frozen policy. (c) Paired differences at 400 ms, where positive values favour the first policy named. The intervals show a clear k=0 advantage over feed-forward and finite-stack memory, but no clear difference from GRU-64 or from adding rank- 1 , rank- 2 , or rank- 4 nonlinear recurrent corrections.
Family
400 ms save (%, 95% CI)
Overall save (%)
Total params
Recurrent core params
Carried state (B)
Recurrent state (B)
MACs per step
CPU latency median / p95 ( μ s)
Feed-forward
44.4 [33.5,56.4]
73.6
5,602
0
0
0
5,440
12.97 / 13.08
Ten-step stack
86.0 [79.0,91.8]
93.7
7,330
0
120
0
7,168
17.86 / 17.99
Structured k=0
98.3 [96.5,99.6]
99.0
12,002
2,304
264
256
11,776
23.03 / 23.15
Structured k=1
97.9 [95.7,99.4]
98.6
12,165
2,467
264
256
11,938
28.20 / 28.39
Structured k=2
98.8 [97.4,99.7]
99.0
12,328
2,630
264
256
12,100
28.26 / 28.43
Structured k=4
98.1 [96.1,99.6]
98.7
12,654
2,956
264
256
12,424
28.31 / 28.61
Table 1: Complete comparison of the seven student families. The 400 ms column reports the five-seed mean and paired-bootstrap 95% confidence interval. Overall save rate pools the five predeclared 0–400 ms conditions and excludes the 500 ms extrapolation. For each latency statistic (median or p95), entries are medians across five seeds of that statistic computed from block-mean latencies on one pinned Intel Core Ultra 9 285 CPU core. Carry includes every value retained between control steps; recurrent state excludes the previous action and the finite observation stack.
Figure 4: Performance–cost frontier at 400 ms tracking loss. Small symbols show the five matched training seeds; large, directly labelled symbols show each family’s five-seed mean save rate. Their horizontal positions use the median measured latency in (a) and the fixed parameter count in (b). Seed points in (b) are offset slightly because all seeds within a family have the same parameter count. The pale dashed line joins the non-dominated observed family summaries. Structured k=0 records 98.3% saves with 12,002 parameters and 23.0 μ s latency, while GRU-64 records 97.8% with 28,898 parameters and 40.1 μ s. Their paired save-rate difference is 0.5 points with a 95% interval of [−1.2,2.5] ; k=0 uses 58.5% fewer parameters and has 42.5% lower measured latency. The figure describes the measured trade-off rather than selecting a universally best architecture.
No blackout
400 ms blackout
Policy / σ
0
1
5
0
1
5
k=0
98.6
98.3
98.4
98.1
98.0
98.5
k=4
98.5
98.6
98.9
97.6
97.5
94.5
GRU-64
99.4
99.2
97.9
96.8
96.8
92.9
Teacher
99.1
99.1
98.7
97.3
96.9
96.0
Table 2: Save rates (%) on 225 fresh shots per condition, without retraining. Student entries average five seeds; the teacher is one frozen policy. Noise σ is per-coordinate Gaussian standard deviation in mm. Paired pointwise 95% intervals for the main differences are reported in the text.
The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning. We provide a theoretical justification for their empirical effectiveness by constructing and studying two linear filters: (i) the first exactly reproduces the pre-softmax logits of the belief vector in a hidden Markov model (HMM) under a deterministic transition matrix, thereby serving as a sufficient statistic for optimal policy learning, (ii) the second achieves vanishing state-decoding error under a nearly deterministic transition matrix, thus reducing state ambiguity to near zero. The results extend to action-controlled HMMs, where the corresponding linear filters become time-varying with action-dependent dynamics. We illustrate our main results through numerical experiments and further show that the constructed linear filter serves as a strong feature extractor in a small reinforcement learning game.
Yike Zhao, Onno Eberhard, Malek Khammassi +2
EPFL, Lausanne, Switzerland · Max Planck Institute for Intelligent Systems, Tübingen, Germany
The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretically realize a set of symbolic transition rules. We argue that equally important is error control, the dynamics governing hidden-state drift along the directions that distinguish symbolic states. We prove that affine recurrent networks, a class of models encompassing State-Space Models and Linear Attention, cannot correct errors along state-separating subspaces once they preserve state representations. Consequently, practical affine trackers do not learn robust state tracking; rather, they learn finite horizon solutions governed by accumulated state-relevant error. We characterize the mechanics of this failure, showing that tracking remains readable only while the accumulating within-class spread remains small relative to the initial between-class separation. We demonstrate empirically on group state-tracking tasks that this breakdown is predictable: tracking collapses when the distinguishability ratio crosses the readability threshold of the trained decoder. Across trained models, the point of this crossing predicts the horizon at which downstream accuracy fails. These results establish that robust state tracking is determined not only by an architecture's theoretical expressivity but crucially by its error control.
Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substantially increase inference cost, or rely on pre-defined semantic features, which limit task generality and may also require the retraining of the backbone to adapt to the memory. We introduce DRAM (Delta-rule Recurrent Associative Memory), a plug-and-play memory module that can be attached to a wide range of pretrained robotic policies, endowing them with long-horizon memory without architectural modification or backbone retraining, requiring only task-specific post-training of the memory module and action expert. DRAM maintains a fixed-size associative memory using gated delta-rule linear attention, with a modified update that incorporates all tokens within each frame in parallel. An architecture-agnostic readout integrates historical context into action prediction across different policy architectures. Experiments show that DRAM consistently improves frozen pretrained policies over short-context baselines and alternative compact memory designs, validating its effectiveness as a fixed-size, post-hoc memory module trained with the backbone frozen.