cs.ROOct 4, 2026

EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation

Authors: Yuheng Na, Zhide Zhong, Junjie He, Junfeng Li, Haodong Yan, Jiaan Wang, Jiaguan Zhu, Yangyang Zheng, +2 more

Organizations: Xi’an Jiaotong University · The Hong Kong University of Science and Technology (Guangzhou) · Xbotics Embodied AI Community · The Chinese University of Hong Kong

Abstract

Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7% on RMBench, 82.0% on RoboMME and 83.8% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

    Sep 29, 2026Yaxin Zhao, Dianye Huang, Chenwei Wang +2Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  2. EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

    Jun 18, 2026Ganlin Yang, Zhangzheng Tu, Yuqiang Yang +10Long-Horizon Robotic ManipulationVision-Language-Action Models

  3. NativeMEM: Native Memory Compression for Long-Horizon Robotic Manipulation

    Jul 7, 2026Ziye Wang, Modi Shi, Chaojun Ni +5Memory-Augmented VLMsLong-Horizon Robotic Manipulation