cs.ROOct 1, 2026

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

Authors: Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh

Organizations: National University of Singapore · Nanyang Technological University, Singapore · LISTENAI

Abstract

Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information I(at;mt∣ot)I(a_t; m_t \mid o_t) between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function mt=M(ht)m_t = M(h_t) and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-KK selection over 2K2K tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

    Sep 29, 2026Yaxin Zhao, Dianye Huang, Chenwei Wang +2Mamba

  2. Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    Jul 8, 2026Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6Generalist Vision--Language--ActionVision-Language-Action Framework

  3. MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation

    Sep 30, 2026Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov +1Goal-Conditioned Dynamic ManipulationRobocasa