cs.ROSep 28, 2026

Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs

Authors: Tanguy Dieudonné, Jack B. Jedlicki, Heng Yang

Organizations: Harvard University · ETH Zürich

Abstract

Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned π0.5π_{0.5} policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

    Sep 29, 2026Yaxin Zhao, Dianye Huang, Chenwei Wang +2Mamba

  2. Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    Jul 8, 2026Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6Generalist Vision--Language--ActionVision-Language-Action Framework

  3. MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation

    Sep 30, 2026Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov +1Goal-Conditioned Dynamic ManipulationRobocasa