Memory-Augmented VLMs

VLM: Vision-Language Model

Latest papers 90

All topics
CardsList
  1. MarvisNav: Making Memory Visible on Route Choices for Zero-Shot Object Navigation

    Oct 5, 2026Jincheng Wang, Chi Pui Chan, Wei Zeng +3Memory-Augmented VLMsVision-Language Navigation

  2. ReMem: Streaming Video Understanding With Long Context Retention

    Oct 5, 2026Li Yiheng, He Xu, Wang Shaobo +2Memory-Augmented VLMsStreaming Video Understanding

  3. EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation

    Oct 4, 2026Yuheng Na, Zhide Zhong, Junjie He +7Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  4. ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory

    Oct 4, 2026Hao Jiang, Zhanyu Guo, Chenwei Wu +11Memory-Augmented VLMsEfficient Multimodal Inference

  5. LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation

    Oct 4, 2026Yiming Zhao, Tianshun Li, Jingle He +2Memory-Augmented VLMsVision-Language Models

  6. VISTA: A Visual Harness for Reasoning in an Interactive World

    Oct 1, 2026Qiushi Han, Keya Hu, Linlu Qiu +2Memory-Augmented VLMsVisual Reasoning

  7. Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

    Oct 1, 2026Xuehui Yu, Eason Yu, Meiyi Wang +3Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  8. MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation

    Sep 30, 2026Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov +1Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  9. Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model

    Sep 30, 2026Zaijing Li, Rui Shao, Bing Hu +3Memory-Augmented VLMsRobot Skill Learning

  10. MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding

    Sep 30, 2026Yinying Li, Yuqian Fu, Yulin Dai +3Memory-Augmented VLMsStreaming Video Understanding

  11. Vision-Language-Action Autonomous Driving Agent with Language-based Memory

    Sep 29, 2026Kai Yan, Xiangyu Chen, Yulong Cao +9Memory-Augmented VLMsVLMs for Autonomous Driving

  12. VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

    Sep 29, 2026Jianming Xu, Jinfa Huang, Jingyang Lin +2Memory-Augmented VLMsLong-Video Understanding

  13. Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

    Sep 29, 2026Yaxin Zhao, Dianye Huang, Chenwei Wang +2Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  14. MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms

    Sep 29, 2026Guohong Liu, Jialei Ye, Shanhui Zhao +2Memory-Augmented VLMsStreaming Video Understanding

  15. D2^2-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

    Sep 28, 2026Zijian Ye, Chengqi Wei, Wei Huang +9Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  16. Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs

    Sep 28, 2026Tanguy Dieudonné, Jack B. Jedlicki, Heng YangMemory-Augmented VLMsLong-Horizon Robotic Manipulation

  17. ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation

    Sep 21, 2026Yibo Li, Enshen Zhou, Rui Chen +7Memory-Augmented VLMsActive Perception

  18. PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

    Sep 20, 2026Siru Zhong, Qiongyan Wang, Xiaohui Lv +5Memory-Augmented VLMsRecurrent Neural Networks

  19. TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

    Sep 20, 2026Hengyan Liu, Wenlve Zhou, Bo Yue +7Memory-Augmented VLMsLong-Horizon Robotic Manipulation

  20. Retrieval Geometry Shapes Cache-Based Clip Adaptation

    Sep 20, 2026Mahir Shahriar Tamim, Md. Samiul Alim, Azmine Toushik Wasi +5Memory-Augmented VLMsTest-Time Adaptation

  21. FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

    Sep 16, 2026Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao +1Memory-Augmented VLMsVLMs for Autonomous Driving

  22. Collaborative Memory for Multi-Agent VLM Systems

    Sep 15, 2026Huixin Zhang, Shao-Jun Xia, Di Wang +3Memory-Augmented VLMsMulti-Agent Collaboration

  23. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

    Sep 15, 2026Dingli Liang, Yiqiao Xie, Yukai Huang +6Memory-Augmented VLMsEgocentric Video QA

  24. AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation

    Sep 14, 2026Shengjie Jin, Zelong Sun, Hengbo Xu +2Memory-Augmented VLMsContinual Learning for VLMs