cs.CVSep 29, 2026

Vision-Language-Action Autonomous Driving Agent with Language-based Memory

Authors: Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, +4 more

Organizations: University of Illinois Urbana-Champaign · NVIDIA

Abstract

Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

    Aug 11, 2026Zebin Xing, Yupeng Zheng, Qiang Chen +10Autonomous DrivingDrives

  2. FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

    Sep 16, 2026Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao +1Diffusion-Based Vision-Language-ActionsAutonomous Driving

  3. Less Language, More Latents: Annotation-Efficient VLAs for Driving

    Sep 23, 2026Alexey Zakharov, Kemal Oksuz, Puneet K. DokaniaDiffusion-Based Vision-Language-ActionsLatent Action Models