Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substantially increase inference cost, or rely on pre-defined semantic features, which limit task generality and may also require the retraining of the backbone to adapt to the memory. We introduce DRAM (Delta-rule Recurrent Associative Memory), a plug-and-play memory module that can be attached to a wide range of pretrained robotic policies, endowing them with long-horizon memory without architectural modification or backbone retraining, requiring only task-specific post-training of the memory module and action expert. DRAM maintains a fixed-size associative memory using gated delta-rule linear attention, with a modified update that incorporates all tokens within each frame in parallel. An architecture-agnostic readout integrates historical context into action prediction across different policy architectures. Experiments show that DRAM consistently improves frozen pretrained policies over short-context baselines and alternative compact memory designs, validating its effectiveness as a fixed-size, post-hoc memory module trained with the backbone frozen.
Figures & tables
Figure 1: Overview of DRAM. Left: DRAM augments representations Xt from a frozen VLA encoder with historical information, producing Yt for action prediction. Right: The memory module follows a read–fuse–write procedure. After reading, current-frame associations update the retained memory through a frame-level delta update, yielding St for the next frame.
Environment
MemoryVLA
HAMLET
μ VLA
ContextVLA π0
ContextVLA π0-FAST
ContextVLA G
DRAM (Ours)
LIBERO-Spatial
98.40
99.00
93.00
97.40
98.30
98.40
99.00
LIBERO-Object
98.40
100.00
99.40
98.20
99.20
99.00
100.00
LIBERO-Goal
96.40
99.20
96.60
96.40
95.60
97.20
98.00
LIBERO-10
93.40
92.20
95.80
93.80
90.20
93.40
99.00
Average
96.65
97.60
96.20
96.45
95.83
97.00
99.00
Table 1: Success rates (%) on LIBERO. Bold indicates the best result in each row, including ties. Superscripts indicate the backbone; DRAM uses π0.5 . Averages are recomputed over the four suites. Published results use their respective training and evaluation protocols.
DiT
π0.5
Suite
Baseline
+ DRAM
Baseline
+ DRAM
LIBERO-Spatial
83.25
92.00
97.00
99.00
LIBERO-Object
88.00
97.50
99.00
100.00
LIBERO-Goal
91.00
98.00
98.00
98.00
LIBERO-10
75.00
91.00
96.00
99.00
Average
84.31
94.63
97.50
99.00
Table 2: LIBERO success rates (%) across policy backbones. Bold indicates the better result within each pair, including ties.
Task
π0.5
X-VLA
Mem-0
DRAM π0.5 (Ours)
Battery Try
16.00
26.00
28.00
65.00
Blocks Ranking Try
6.00
1.00
18.00
45.00
Cover Blocks
0.00
2.00
68.00
70.00
Observe and Pick Up
9.00
9.00
4.00
15.00
Press Button
0.00
0.00
0.00
20.00
Put Back Block
11.00
18.00
90.00
65.00
Table 3: Success rates (%) on RMBench. Bold indicates the best result in each row, including ties. π0.5 , X-VLA, and Mem-0 results are taken from the RMBench benchmark and use their respective training and evaluation protocols; DRAM uses the π0.5 backbone with native-KV integration.
Baseline
DRAM (Ours)
Suite
π0.5
Native KV
Hidden Tokens
Last Hidden
LIBERO-Spatial
97.00
99.00
98.50
98.50
LIBERO-Object
99.00
100.00
99.00
99.50
LIBERO-Goal
98.00
98.00
99.50
96.00
LIBERO-10
96.00
99.00
97.50
96.00
Average
97.50
99.00
98.63
97.50
Table 4: Success rates (%) of different DRAM configurations on LIBERO with π0.5 from LeRobot. Averages are recomputed over the four displayed suite-level success rates. Bold indicates the best result in each row.
Figure 2: LIBERO-10 success versus memory reset interval L (in frames). Full resets only between episodes; the dashed line denotes the π0.5 baseline without DRAM.
Task
DiT
DiT + DRAM
Insertion
90.0
100.0
Full Task
0.0
85.0
Table 5: Real-world success rates (%).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Write-source configurations for the π0.5 host (separate-source integration). DRAM layers are aligned with the VLM backbone layers, and the write sources of memory layer ℓ are tapped from (i) the backbone attention’s native key–value projections ( Native KV ), (ii) intermediate hidden tokens ( Hidden tokens ), or (iii) only the final-layer hidden states ( Last hidden ). Query slots, readout injection, and the memory update are identical across configurations.
Figure 4: Real-world experimental platform and manipulation tasks. (a) The I7 Pro robot. (b) The Install task, which requires precise alignment and insertion from a predefined pose near the target. (c) The Full task, which extends Install with an initial approach, workpiece release, and arm retraction to complete the loading sequence. For each task, three representative execution stages are illustrated using a global view, a close-up view of the manipulation, and a wrist-camera observation.
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T2Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T2Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T2Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/
Yize Liu, Huang Huang, Yining Hong +4
Stanford University · NVIDIA · University of Michigan, Ann Arbor
Long-horizon robot manipulation requires policies to track completed subtasks and critical interaction events. However, existing memory mechanisms heavily rely on external models or predefined update rules. To address this, we propose OnEvoMemory, a value-guided memory module for pretrained robot policies. It maintains recent context, high-value experiences, and salient transitions, while learning which experiences should be retained from trajectory outcomes. Offline demonstrations initialize the memory prior, whereas successful and unsuccessful online rollouts refine memory selection, helping the policy recognize task-stage transitions and avoid repeating completed subtasks. Experiments on long-horizon manipulation benchmarks show that OnEvoMemory improves the performance of the base VLA policy through both offline initialization and online memory evolution.
Zhongxi Chen, Shenqi Zong
Shanghai Jiao Tong University · Tsinghua University
Many robotic tasks require short-term memory, whether it's retrieving an object that's no longer visible or turning off an appliance after a set period. Yet, most visuomotor policies trained via imitation learning rely only on immediate sensory input without using past experiences to guide decisions. We present PRISM, a transformer-based architecture for visuomotor policies to effectively use short-term memory via two key components: (i) gated attention, which filters retrieved information to suppress irrelevant details, improving performance by reducing the spurious correlations between the history and current action prediction, (ii) a hierarchical architecture that first compresses local information into compact tokens and then integrates them to capture temporally extended dependencies, improving its compute and memory footprint. Together, these mechanisms enable us to scale short-term memory in visuomotor policies for up to two minutes. To systematically evaluate memory in visuomotor control, we introduce ReMemBench -- a benchmark of eight diverse household manipulation tasks spanning four categories of short-term memory -- designed to foster general memory mechanisms rather than siloed, task-specific solutions. PRISM consistently outperforms prior works, including recurrent architectures, transformers, and their variants -- achieving an absolute improvement of 5%--12% over the strongest baseline. On the RoboCasa and LIBERO benchmarks, it achieves absolute improvements of 11%--15% over its no-memory variant and fine-tuned Vision-Language-Action baselines such as GR00T-N1-3B and OpenVLA, despite not leveraging any large-scale pretraining. Together, PRISM and ReMemBench establish a foundation for developing and evaluating short-term memory-augmented visuomotor policies that scale to long-horizon tasks. Additional materials are available at https://shahrutav.github.io/short-term-memory
Rutav Shah, Rajat Kumar Jenamani, Xiaohan Zhang +5
1Robotics and AI Institute, 2The University of Texas at Austin