Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 × the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: https://github.com/Sathya98/qurl
Figures & tables
Figure 2: Reward curves for the 3 uRNN variants and the baselines across 7 environments from the POBAX ( Tao et al., 2025 ) benchmark
Figure 3: Reward curves for the 3 Phase Aware QuRNN variants plotted with a GRU baseline and the best uRNN method on 5 environments with discrete action spaces
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Geometric interpretation of a Householder reflection. A vector x is reflected across the hyperplane orthogonal to v to produce Hx , where H=I−2vv∗/(v∗v) is the Householder matrix. The reflection flips the parallel component while preserving the perpendicular component.
Hyperparameter
Value
Note
Optimizer
Adam
Update epochs
4
per PPO update
Minibatches per update
4
PPO clip ϵ
0.2
Value loss coefficient cv
0.5
Max gradient norm
0.5
global-norm clipping
Appendix
Table 1: Hyperparameters held fixed across all baselines and all six uRNN variants. The rollout length T is the only PPO setting we change relative to the baselines (see note).
Environment
Hidden size n
Parallel envs
Total steps
γ
T-Maze (corridor 75)
32
4
5.0×106
0.99
RockSample (11,11)
256
8
5.0×106
0.99
RockSample (15,15)
512
16
1.0×107
0.999
Battleship (10×10)
512
32
1.0×107
1.0
Masked Walker
256
4
5.0×107
0.99
Masked HalfCheetah
256
4
5.0×107
0.99
Appendix
Table 2: Per-environment settings, fixed across all methods. The complex hidden state of each uRNN variant has the same dimension n as the GRU baseline’s hidden state. Training budgets follow the POBAX benchmark [ Tao et al., 2025 ] ; Craftax is capped at 108 steps.
GRU-PPO
PPO-LD
Environment
lr
ent.
λ0
lr
ent.
λ0
λ1
β
T-Maze (corridor 75)
2.5×10−3
0.01
0.7
2.5×10−4
0.01
0.95
0.95
0.25
RockSample (11,11)
2.5×10−3
0.2
0.7
2.5×10−3
0.2
0.5
0.5
0.25
RockSample (15,15)
2.5×10−3
0.2
0.7
2.5×10−3
0.2
0.1
0.95
0.5
Battleship (10×10)
2.5×10−3
0.05
0.7
2.5×10−3
0.05
0.1
0.95
0.5
Masked Walker
2.5×10−4
0.01
0.95
2.5×10−4
0.01
0.95
0.95
0.5
Appendix
Table 3: Per-environment best hyperparameters for the baselines, selected by a grid sweep, exactly following [ Tao et al., 2025 ]
Parameter
Values swept
complex lr
10−6,5×10−6,10−5,
5×10−5,10−4
entropy
0.005,0.01,0.05
λ0
0.8,0.9,0.95
lr
2.5×10−4,2.5×10−3
Appendix
Table 4: T-Maze (corridor 75). Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.075,0.1,0.15,0.2
λ0
0.7,0.8,0.95
lr
2.5×10−4,2.5×10−3
Appendix
Table 5: RockSample (11,11) . Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.075,0.1,0.15,0.2
λ0
0.7,0.8,0.95
lr
2.5×10−4,2.5×10−3
Appendix
Table 6: RockSample (15,15) . Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.01,0.05,0.1
λ0
0.6,0.7,0.8,0.9,0.95
lr
2.5×10−3
Appendix
Table 7: Battleship (10×10) . Search space (left) and best per-variant settings (right).
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.005,0.01,0.05,0.1
λ0
0.8,0.9,0.95
lr
2.5×10−4
Appendix
Table 8: Masked Walker. Search space (left) and best per-variant settings (right). The Born-rule (Q) variants are defined for discrete action spaces only and are not evaluated here.
Parameter
Values swept
complex lr
10−6,10−5,8×10−5
entropy
0.005,0.01,0.05,0.1
λ0
0.8,0.9,0.95
lr
2.5×10−4
Appendix
Table 9: Masked HalfCheetah. Search space (left) and best per-variant settings (right). The Born-rule (Q) variants are defined for discrete action spaces only and are not evaluated here.