Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 × the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: https://github.com/Sathya98/qurl
Figures & tables
Figure 2: Reward curves for the 3 uRNN variants and the baselines across 7 environments from the POBAX ( Tao et al., 2025 ) benchmark
Figure 3: Reward curves for the 3 Phase Aware QuRNN variants plotted with a GRU baseline and the best uRNN method on 5 environments with discrete action spaces
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Geometric interpretation of a Householder reflection. A vector x is reflected across the hyperplane orthogonal to v to produce Hx , where H=I−2vv∗/(v∗v) is the Householder matrix. The reflection flips the parallel component while preserving the perpendicular component.
Hyperparameter
Value
Note
Optimizer
Adam
Update epochs
4
per PPO update
Minibatches per update
4
PPO clip ϵ
0.2
Value loss coefficient cv
0.5
Max gradient norm
0.5
global-norm clipping
Appendix
Table 1: Hyperparameters held fixed across all baselines and all six uRNN variants. The rollout length T is the only PPO setting we change relative to the baselines (see note).
Environment
Hidden size n
Parallel envs
Total steps
γ
T-Maze (corridor 75)
32
4
5.0×106
0.99
RockSample (11,11)
256
8
5.0×106
0.99
RockSample (15,15)
512
16
1.0×107
0.999
Battleship (10×10)
512
32
1.0×107
1.0
Masked Walker
256
4
5.0×107
0.99
Masked HalfCheetah
256
4
5.0×107
0.99
Appendix
Table 2: Per-environment settings, fixed across all methods. The complex hidden state of each uRNN variant has the same dimension n as the GRU baseline’s hidden state. Training budgets follow the POBAX benchmark [ Tao et al., 2025 ] ; Craftax is capped at 108 steps.
GRU-PPO
PPO-LD
Environment
lr
ent.
λ0
lr
ent.
λ0
λ1
β
T-Maze (corridor 75)
2.5×10−3
0.01
0.7
2.5×10−4
0.01
0.95
0.95
0.25
RockSample (11,11)
2.5×10−3
0.2
0.7
2.5×10−3
0.2
0.5
0.5
0.25
RockSample (15,15)
2.5×10−3
0.2
0.7
2.5×10−3
0.2
0.1
0.95
0.5
Battleship (10×10)
2.5×10−3
0.05
0.7
2.5×10−3
0.05
0.1
0.95
0.5
Masked Walker
2.5×10−4
0.01
0.95
2.5×10−4
0.01
0.95
0.95
0.5
Appendix
Table 3: Per-environment best hyperparameters for the baselines, selected by a grid sweep, exactly following [ Tao et al., 2025 ]
Parameter
Values swept
complex lr
10−6,5×10−6,10−5,
5×10−5,10−4
entropy
0.005,0.01,0.05
λ0
0.8,0.9,0.95
lr
2.5×10−4,2.5×10−3
Appendix
Table 4: T-Maze (corridor 75). Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.075,0.1,0.15,0.2
λ0
0.7,0.8,0.95
lr
2.5×10−4,2.5×10−3
Appendix
Table 5: RockSample (11,11) . Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.075,0.1,0.15,0.2
λ0
0.7,0.8,0.95
lr
2.5×10−4,2.5×10−3
Appendix
Table 6: RockSample (15,15) . Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.01,0.05,0.1
λ0
0.6,0.7,0.8,0.9,0.95
lr
2.5×10−3
Appendix
Table 7: Battleship (10×10) . Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.005,0.01,0.05,0.1
λ0
0.8,0.9,0.95
lr
2.5×10−4
Appendix
Table 8: Masked Walker. Search space (left) and best per-variant settings (right). The Born-rule (Q) variants are defined for discrete action spaces only and are not evaluated here.
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.005,0.01,0.05,0.1
λ0
0.8,0.9,0.95
lr
2.5×10−4
Appendix
Table 9: Masked HalfCheetah. Search space (left) and best per-variant settings (right). The Born-rule (Q) variants are defined for discrete action spaces only and are not evaluated here.
Streaming reinforcement learning has emerged as an online learning paradigm that conforms to the restrictions of natural learning agents that process data incrementally, i.e. with a batch size of 1 and no replay buffer. While streaming RL has recently been shown to scale with deep function approximation with full observability, partially observable settings have remained out of reach. Truncated backpropagation through time collapses to a one-step gradient horizon under the streaming setting, and exact real-time recurrent learning is prohibitively expensive. We close this gap using recurrent trace units, a diagonal recurrent architecture that enables exact RTRL with linear time and memory complexity in the parameter count, and show that they integrate cleanly into existing streaming algorithms across both discrete and continuous control. On a MemoryChain diagnostic with chain lengths from 2 to 128, our method sustains performance where streaming TBPTT(1) baselines using feedforward, GRU, and RTU networks collapse. On five POPGym tasks and on partially observable MuJoCo continuous control, the streaming approach is competitive with batched PPO on POPGym and recovers a substantial fraction of batched performance on masked MuJoCo, despite using no replay buffer or batched updates.
Noah Farr, Aryaman Reddi, Carlo D'Eramo +1
Technical University of Darmstadt · Zuse School ELIZA · Center for Artificial Intelligence and Data Science, University of Würzburg +1
In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least 0.82 in all sixteen Endless T-Maze configurations and at least 0.99 on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: https://quartz-admirer.github.io/ALER-Adaptive-Learnable-Experience-Rewriting/.
Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov +1
Innopolis University, Innopolis, Russia · MIRIAI, Moscow, Russia · Cognitive AI Systems Lab, Moscow, Russia
The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning. We provide a theoretical justification for their empirical effectiveness by constructing and studying two linear filters: (i) the first exactly reproduces the pre-softmax logits of the belief vector in a hidden Markov model (HMM) under a deterministic transition matrix, thereby serving as a sufficient statistic for optimal policy learning, (ii) the second achieves vanishing state-decoding error under a nearly deterministic transition matrix, thus reducing state ambiguity to near zero. The results extend to action-controlled HMMs, where the corresponding linear filters become time-varying with action-dependent dynamics. We illustrate our main results through numerical experiments and further show that the constructed linear filter serves as a strong feature extractor in a small reinforcement learning game.
Yike Zhao, Onno Eberhard, Malek Khammassi +2
EPFL, Lausanne, Switzerland · Max Planck Institute for Intelligent Systems, Tübingen, Germany