cs.LGSep 29, 2026

Reinforcement Learning with Complex (valued) Memories

Authors: Sathya Kamesh Bhethanabhotla, Efstratios Gavves, André Biedenkapp

Organizations: University of Freiburg

Abstract

Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 ×\times the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: https://github.com/Sathya98/qurl

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning

    May 23, 2026Noah Farr, Aryaman Reddi, Carlo D'Eramo +1Recurrent Neural NetworksDeterministic Replay

  2. ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning

    Sep 30, 2026Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov +1Prioritized Experience ReplayOffline Reinforcement Learning

  3. Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning

    May 29, 2026Yike Zhao, Onno Eberhard, Malek Khammassi +2Recurrent ModelPartially Observable Markov Decision Process