In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least 0.82 in all sixteen Endless T-Maze configurations and at least 0.99 on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: https://quartz-admirer.github.io/ALER-Adaptive-Learnable-Experience-Rewriting/.
Figures & tables
Figure 1: ALER overview. The recurrent state ht queries a slot memory, a learned gate fuses the retrieved vector rt with ht into zt for the actor and critic, and an independently addressed write stores the write candidate xt .
Observation
Effect
Corridor step, NoOp
none
New cue (Endless)
y←c′
Invert
y←K−1−y
SkipNext
cancels next rune
ResetRules
y←c , flags cleared
RepeatPrev
repeats last rune
Table 1: Update effects and memory counts of the Rune-Mazes tasks. Left: effect of each observation on the target y and the rule state. Right: for each rune composition, the number of rune states Q and the number of target states Q⋆ over all rune sequences ( Proposition 2 ), and the exact counts Qep and Qep⋆ for sampled episodes ( Appendix A ). The deferred solution uses KQ states. Rune Multi-Corridor has K branches.
Figure 2: ALER architecture. The shared LSTM supplies the read query, the recurrent input to fusion, and the write candidate. Reading uses Mt−1 , and independent writes produce Mt .
Figure 3: Rune-Mazes environments. Top left: Rune Multi-Corridor, in which a multi-bit cue indicates the target branch and an Invert rune before the junction changes it. Top right: Rune T-Maze, in which the agent must invert the initial cue after an Invert rune and ignore a NoOp rune. Bottom: Rune MiniGrid Memory with RepeatPrev + SkipNext + Invert , one episode from the cue (left) to the junction (right). The agent crosses RepeatPrev (red), Invert (purple), and SkipNext (yellow), and each rune fires once on contact (legend in Table 7 ).
Figure 4: Endless T-Maze horizon extrapolation. Mean success rate over five training seeds for agents trained on 10 corridors of length 10 and evaluated on 10 corridors of lengths 10 to 100 . Shaded regions show the SEM.
Fixed length
Random length
PPO-LSTM
ALER
PPO-LSTM
ALER
I
0.80±0.03
0.69±0.09
0.93±0.07
0.94±0.06
N+I
0.94±0.06
0.77±0.03
0.67±0.07
0.92±0.08
S+I
0.51±0.03
0.59±0.01
0.62±0.07
0.63±0.05
3I+R
0.53±0.04
0.57±0.01
0.53±0.01
0.57±0.01
P+S+I
0.49±0.03
0.60±0.00
0.56±0.00
0.69±0.06
Table 5: Rune MiniGrid Memory. Success rate of the selected checkpoint within 30 M steps on the 100 validation episodes, mean ± STD over three training seeds. Bold marks the higher mean. Rows are compositions with I = Invert , N = NoOp , S = SkipNext , R = ResetRules , and P = RepeatPrev .
Table 6: Training and evaluation configurations for Rune T-Maze and Rune Multi-Corridor. Each row is a separate training task.
Tile
Rune (color)
Effect
Invert (purple)
Swaps the hidden success and failure targets. The two visible objects stay in place.
SkipNext (yellow)
Cancels the next rune the agent steps on. That rune is consumed without effect.
ResetRules (green)
Restores the original targets, clears the Invert and SkipNext effects, and turns the agent east.
RepeatPrev (red)
Re-applies the last rune that took effect and has no effect if no rune has fired.
NoOp (grey)
Leaves the rules unchanged and appears as a visual distractor.
Appendix
Table 7: Rune legend for Rune MiniGrid Memory. Each rune is a floor tile of its own color.
Figure 5: Example Rune MiniGrid Memory rollouts. From top to bottom: no runes, Invert , NoOp + Invert , SkipNext + Invert , 3 Invert + ResetRules , and RepeatPrev + SkipNext + Invert . Each strip follows one episode from the cue to the junction.
Configuration
PPO-LSTM
ALER
S9
0.66±0.13
0.84±0.22
S9Random
0.83±0.22
1.00±0.00
S11
0.55±0.01
0.67±0.19
S11Random
0.84±0.23
0.95±0.07
S13
0.55±0.00
0.83±0.23
S13Random
0.78±0.16
0.97±0.04
Appendix
Table 9: MiniGrid Memory. Success rate of the selected checkpoint on the 100 validation episodes, mean ± SEM over three training seeds. Bold marks the higher mean in each row.
Figure 6: Rune Multi-Corridor learning curves. Validation return and success rate over training for ALER, S5, PPO-LSTM, GTrXL, FFM, SHM, and PPO-MLP for K∈{4,8,16} branches, with ( rune=1 ) and without ( rune=0 ) an Invert rune.
Figure 7: Rune T-Maze learning curves. Validation return and success rate over training for ALER, S5, PPO-LSTM, GTrXL, FFM, SHM, and PPO-MLP. Column titles give the per-type rune counts of each composition.
Gate
ET
RT-1
RT-2
Learned gt
0.72±0.14
1.00±0.00
1.00±0.00
gt=0
0.48±0.29
0.83±0.02
0.72±0.04
gt=0.5
0.06±0.06
0.45±0.11
0.54±0.13
gt=1
0.00±0.00
0.26±0.13
0.21±0.11
Appendix
Table 10: Gate experiments on trained ALER. Success rate, mean ± SEM over three training seeds, on evaluation seeds 1 – 3 with 100 episodes each. ET is Endless T-Maze with 20 corridors of length 50 , and RT-1 and RT-2 are Rune T-Maze Invert and RepeatPrev + SkipNext + Invert .
Figure 8: Cue decoding at Endless T-Maze junctions. Current- and previous-corridor cue decoding from eight representations with 20 corridors of the indicated length. Each probe is fitted and tested at that length on disjoint sets of episodes. Cells report balanced accuracy, mean ± SEM over three training seeds.
Figure 9: Current-cue decoding within Endless T-Maze corridors. Corridors have length 50 , with 20 corridors per episode. Only the first observation at each sampled position is kept, and the cue is visible at position 0 and masked at positions 12 , 25 , 38 , and 50 . Left: probes fitted at each position. Right: probes fitted at position 0 and applied without refitting. Mt− precedes the write and Mt+ includes it. Cells report balanced accuracy, mean ± SEM over three training seeds.
Figure 10: Transfer of Endless T-Maze probes fitted at length 20. Length 20 uses held-out episodes, and longer lengths use independent sets of 128 episodes per checkpoint. Points and error bars show the mean balanced accuracy and SEM over three training seeds, with small horizontal offsets between overlapping estimates.
Figure 11: Probes on Rune T-Maze. Corridor length 10 and rune probability 0.80 . All variables are decoded at junctions. The last column uses probes fitted and tested only on successful episodes, and the other columns use all episodes. Cells report balanced accuracy, mean ± SEM over three training seeds.
Figure 12: Gate dynamics in Rune T-Maze. Top: mean gate coefficient aligned to rune observations, with the rune step highlighted. Bottom: mean absolute gate-component change at rune steps and at matched non-rune steps. Shaded bands and error bars show the SEM over three training seeds.