Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substantially increase inference cost, or rely on pre-defined semantic features, which limit task generality and may also require the retraining of the backbone to adapt to the memory. We introduce DRAM (Delta-rule Recurrent Associative Memory), a plug-and-play memory module that can be attached to a wide range of pretrained robotic policies, endowing them with long-horizon memory without architectural modification or backbone retraining, requiring only task-specific post-training of the memory module and action expert. DRAM maintains a fixed-size associative memory using gated delta-rule linear attention, with a modified update that incorporates all tokens within each frame in parallel. An architecture-agnostic readout integrates historical context into action prediction across different policy architectures. Experiments show that DRAM consistently improves frozen pretrained policies over short-context baselines and alternative compact memory designs, validating its effectiveness as a fixed-size, post-hoc memory module trained with the backbone frozen.
Figures & tables
Figure 1: Overview of DRAM. Left: DRAM augments representations Xt from a frozen VLA encoder with historical information, producing Yt for action prediction. Right: The memory module follows a read–fuse–write procedure. After reading, current-frame associations update the retained memory through a frame-level delta update, yielding St for the next frame.
Environment
MemoryVLA
HAMLET
μ VLA
ContextVLA π0
ContextVLA π0-FAST
ContextVLA G
DRAM (Ours)
LIBERO-Spatial
98.40
99.00
93.00
97.40
98.30
98.40
99.00
LIBERO-Object
98.40
100.00
99.40
98.20
99.20
99.00
100.00
LIBERO-Goal
96.40
99.20
96.60
96.40
95.60
97.20
98.00
LIBERO-10
93.40
92.20
95.80
93.80
90.20
93.40
99.00
Average
96.65
97.60
96.20
96.45
95.83
97.00
99.00
Table 1: Success rates (%) on LIBERO. Bold indicates the best result in each row, including ties. Superscripts indicate the backbone; DRAM uses π0.5 . Averages are recomputed over the four suites. Published results use their respective training and evaluation protocols.
DiT
π0.5
Suite
Baseline
+ DRAM
Baseline
+ DRAM
LIBERO-Spatial
83.25
92.00
97.00
99.00
LIBERO-Object
88.00
97.50
99.00
100.00
LIBERO-Goal
91.00
98.00
98.00
98.00
LIBERO-10
75.00
91.00
96.00
99.00
Average
84.31
94.63
97.50
99.00
Table 2: LIBERO success rates (%) across policy backbones. Bold indicates the better result within each pair, including ties.
Task
π0.5
X-VLA
Mem-0
DRAM π0.5 (Ours)
Battery Try
16.00
26.00
28.00
65.00
Blocks Ranking Try
6.00
1.00
18.00
45.00
Cover Blocks
0.00
2.00
68.00
70.00
Observe and Pick Up
9.00
9.00
4.00
15.00
Press Button
0.00
0.00
0.00
20.00
Put Back Block
11.00
18.00
90.00
65.00
Table 3: Success rates (%) on RMBench. Bold indicates the best result in each row, including ties. π0.5 , X-VLA, and Mem-0 results are taken from the RMBench benchmark and use their respective training and evaluation protocols; DRAM uses the π0.5 backbone with native-KV integration.
Baseline
DRAM (Ours)
Suite
π0.5
Native KV
Hidden Tokens
Last Hidden
LIBERO-Spatial
97.00
99.00
98.50
98.50
LIBERO-Object
99.00
100.00
99.00
99.50
LIBERO-Goal
98.00
98.00
99.50
96.00
LIBERO-10
96.00
99.00
97.50
96.00
Average
97.50
99.00
98.63
97.50
Table 4: Success rates (%) of different DRAM configurations on LIBERO with π0.5 from LeRobot. Averages are recomputed over the four displayed suite-level success rates. Bold indicates the best result in each row.
Figure 2: LIBERO-10 success versus memory reset interval L (in frames). Full resets only between episodes; the dashed line denotes the π0.5 baseline without DRAM.
Task
DiT
DiT + DRAM
Insertion
90.0
100.0
Full Task
0.0
85.0
Table 5: Real-world success rates (%).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Write-source configurations for the π0.5 host (separate-source integration). DRAM layers are aligned with the VLM backbone layers, and the write sources of memory layer ℓ are tapped from (i) the backbone attention’s native key–value projections ( Native KV ), (ii) intermediate hidden tokens ( Hidden tokens ), or (iii) only the final-layer hidden states ( Last hidden ). Query slots, readout injection, and the memory update are identical across configurations.
Figure 4: Real-world experimental platform and manipulation tasks. (a) The I7 Pro robot. (b) The Install task, which requires precise alignment and insertion from a predefined pose near the target. (c) The Full task, which extends Install with an initial approach, workpiece release, and arm retraction to complete the loading sequence. For each task, three representative execution stages are illustrated using a global view, a close-up view of the manipulation, and a wrist-camera observation.