cs.CVSep 28, 2026

D2^2-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

Authors: Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, +4 more

Organizations: The University of Hong Kong · Southern University of Science and Technology

Abstract

Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D2^2-VLA, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. D2^2-VLA uses block-wise causal KV caching to encode observations incrementally and, guided by distinct temporal attention patterns, constructs separate historical KV read views for the VLM and action expert. Between periodic VLM updates, a gated adapter incorporates fresh visual features into the latest history-conditioned KV block, while a short fast-memory queue supports action replanning. We introduce DOMINO-Long, a ten-task benchmark requiring robots to use earlier visual cues when manipulating moving objects. D2^2-VLA achieves complete-task success rates of 29.3% on DOMINO, compared with 9.6% for π0.5π_{0.5} and 17.2% for PUMA, and 60.0% on DOMINO-Long, compared with 35.4% and 20.6%, respectively. It improves success rates on eight real-robot tasks and reaches 97.5% on LIBERO-Long and 74.3% on RoboTwin 2.0.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

    Sep 29, 2026Yaxin Zhao, Dianye Huang, Chenwei Wang +2Mamba

  2. MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

    Jun 6, 2026Shanglin Yuan, Weiheng Zhao, Xianda Guo +4Long-Horizon ManipulationMotion

  3. Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    Jul 8, 2026Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6Generalist Vision--Language--ActionVision-Language-Action Framework