cs.LGSep 30, 2026

ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning

Authors: Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev

Organizations: Innopolis University, Innopolis, Russia · MIRIAI, Moscow, Russia · Cognitive AI Systems Lab, Moscow, Russia

Abstract

In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least 0.820.82 in all sixteen Endless T-Maze configurations and at least 0.990.99 on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: https://quartz-admirer.github.io/ALER-Adaptive-Learnable-Experience-Rewriting/.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Forager: a lightweight testbed for continual learning with partial observability in RL

    May 1, 2026Steven Tang, Xinze Xiong, Anna Hakhverdyan +7Offline Reinforcement LearningPartial Observability

  2. Replay What Matters: Off-Policy Replay for Efficient LLM Reinforcement Unlearning

    Jun 13, 2026Zirui Pang, Chenlong Zhang, Haosheng Tan +3Large Language Model UnlearningLarge Language Model Reinforcement Learning