cs.ROSep 29, 2026

Simple Agentic Memory for Generalist Robot Policies

Authors: Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin, Shilong Liu, Ding Zhao

Organizations: Carnegie Mellon University · Tulane University · New York University · Princeton University · Columbia University

Abstract

Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ECoMEM: Explicit Concept Memory for Memory-Dependent Robot Control

    Sep 30, 2026Yize Liu, Ke Wang, Mac Schwager +2Robot StateRobot Systems

  2. MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

    Sep 23, 2026Tej Deep Pala, Navonil Majumder, Bryce Goh +4Generalizable Vision-Language-Action PoliciesEmbodied

  3. T2^2Mem: Learning Test-Time Memory for Robotics

    Sep 29, 2026Yize Liu, Huang Huang, Yining Hong +4Robotic ManipulationTest Time