Vision-Language-Action Autonomous Driving Agent with Language-based Memory
Organizations: University of Illinois Urbana-Champaign · NVIDIA
Abstract
Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
Figures & tables
| minADE | Avg. ADE | ML. ADE | Stop SR | Go SR | Roll | MCQ Acc. | |||
| Base Model | 1.386 | 2.351 | 2.393 | 82.06% | 33.74% | 1.725 | 1.871 | 11.79% | 0% |
| Alpamayo 2 Super 34B | 0.993 | 2.293 | N/A | 78.75% | 34.30% | 2.182 | 1.488 | 15.94% | 33.19% |
| No mem. (SFT only) | 1.102 | 2.185 | 2.318 | 87.04% | 38.59% | 1.358 | 1.564 | 6.62% | 33.77% |
| CoT mem. (SFT only) | 1.096 | 2.171 | 2.296 | 86.89% | 39.37% | 1.377 | 1.531 | 6.69% | 45.41% |
| AD-Memo (SFT only) | 1.049 | 2.121 | 2.256 | 87.48% | 40.66% | 1.257 | 1.469 | 6.05% | 45.99% |
| CoT mem. (SFT+ref. mem) | 1.098 | 2.169 | 2.286 | 86.95% | 39.45% | 1.367 | 1.521 | 6.61% | 45.24% |
| minADE | Avg. ADE | ML. ADE | Stop SR | Go SR | Roll | MCQ Acc. | |||
|---|---|---|---|---|---|---|---|---|---|
| Da Capo (Ours) | 0.951 | 1.866 | 1.915 | 89.26% | 45.01% | 1.070 | 1.183 | 5.26% | 51.66% |
| Capo (no std) | 0.942 | 1.952 | 2.047 | 89.22% | 44.27% | 1.130 | 1.277 | 5.01% | 50.91% |
| Capo + const. scaling | 0.943 | 1.911 | 1.979 | 89.21% | 44.74% | 1.122 | 1.227 | 5.03% | 51.57% |
| GRPO (traj-level reward) | 1.007 | 2.113 | 2.194 | 87.11% | 40.94% | 1.312 | 1.392 | 6.28% | 48.16% |
| AD-Memo (SFT only) | 1.049 | 2.121 | 2.256 | 87.48% | 40.66% | 1.257 | 1.469 | 6.05% | 45.99% |
| minADE | Avg. ADE | ML. ADE | minFDE | Avg. FDE | ML. FDE | Corner | MCQ Acc. | |
| Base Model | 1.032 | 1.981 | 2.049 | 2.853 | 5.937 | 6.194 | 1.001 | 0 |
| Alpamayo 2 Super 34B | 0.954 | 2.081 | N/A | 2.556 | 6.141 | N/A | 0.902 | 25.71% |
| No mem. (SFT only) | 1.007 | 2.033 | 2.075 | 2.716 | 6.108 | 6.279 | 0.968 | 58.02% |
| CoT mem. (SFT only) | 1.031 | 2.059 | 2.105 | 2.787 | 6.192 | 6.375 | 0.993 | 57.51% |
| AD-Memo (SFT only) | 1.005 | 2.011 | 2.049 | 2.705 | 6.042 | 6.196 | 0.967 | 65.63% |
| CoT mem. (SFT+ref. mem) | 1.034 | 2.064 | 2.105 | 2.795 | 6.207 | 6.373 | 0.995 | 57.36% |
| Our memory + last frame | Unmerged memory + last frame | Only last frame | Full video | |
|---|---|---|---|---|
| WaymoQA | 71.65% | 68.3% | 66.63% | 75% |
| LingoQA | 67% | 60.4% | 62.6% | 70% |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Supervised Fine-Tuning | |
| global batch size | 128 |
| learning rates | (vision encoder); (language model and LM head) |
| optimizer | fused AdamW |
| Adam betas / epsilon | / |
| weight decay | 0.1 |
| lr scheduler | cosine decay; 100-step linear warm-up; minimum lr factor 0.01 |
| minADE | Avg. ADE | ML. ADE | Stop SR | Go SR | Roll | MCQ Acc. | |||
|---|---|---|---|---|---|---|---|---|---|
| Alpamayo 2 Super 34B | 0.993 | 2.293 | N/A | 78.75% | 34.30% | 2.182 | 1.488 | 15.94% | 33.19% |
| Base model (2B) | 1.386 | 2.351 | 2.393 | 82.06% | 33.74% | 1.725 | 1.871 | 11.79% | 0% |
| No memory (2B SFT) | 1.102 | 2.185 | 2.318 | 87.04% | 38.59% | 1.358 | 1.564 | 6.62% | 33.77% |
| CoT-as-memory (2B SFT) | 1.096 | 2.171 | 2.296 | 86.89% | 39.37% | 1.377 | 1.531 | 6.69% | 45.41% |
| Ours (2B SFT) | 1.049 | 2.121 | 2.256 | 87.48% | 40.66% | 1.257 | 1.469 | 6.05% | 45.99% |
| Base model (8B) | 1.315 | 2.254 | 2.332 | 83.72% | 35.65% | 1.627 | 1.776 | 10.90% | 0% |
| Easy | Hard (#1, past-anchor) | Hard (#2, multi-hop) | Avg. test set | |
| Random | 16.67% | 16.67% | 16.67% | 16.67% |
| Alpamayo 2 Super 34B | 48.45% | 25.77% | 25.65% | 25.71% |
| No memory | 98.14% | 54.93% | 61.10% | 58.02% |
| CoT-as-Memory | 98.02% | 51.46% | 63.56% | 57.51% |
| Ours | 97.92% | 60.92% | 70.34% | 65.63% |
| CoT-as-Memory (GT mem) | 97.99% | 51.25% | 63.48% | 57.36% |
| ADE gain | FDE gain | Shares | |
|---|---|---|---|
| (min/Avg./ML) | (min/Avg./ML) | (%) | |
| Vulnerable Road Users | -0.011 /+0.010/ +0.018 | -0.011 / +0.054 / +0.072 | 41.46% |
| Traffic Control Compliance | +0.013 / +0.023 / +0.030 | +0.089 / +0.101 / +0.105 | 20.57% |
| Lead Vehicle Following | +0.003/ +0.017 / +0.012 | +0.021 / +0.058 / +0.031 | 8.98% |
| Turning Maneuver | -0.002/ +0.024 / +0.031 | -0.019 / +0.061 / +0.092 | 7.81% |
| Stop for Vehicles | +0.008/ +0.011 / +0.017 | +0.057 / +0.072 / +0.067 | 7.79% |
| Model | Decode success rate |
|---|---|
| All-way stop dataset | |
| Base model | 100.00% |
| Alpamayo 2 Super 34B | 100.00% |
| No memory (SFT) | 99.44% |
| CoT-as-Memory (SFT) | 99.41% |
| Ours (SFT) | 99.51% |
| minADE | Avg. ADE | ML. ADE | Stop SR | Go SR | Roll | MCQ Acc. | |||
|---|---|---|---|---|---|---|---|---|---|
| 0.02 | 0.945 | 1.886 | 1.927 | 88.73% | 44.80% | 1.111 | 1.184 | 5.54% | 50.43% |
| 0.05 | 0.943 | 1.877 | 1.913 | 88.67% | 45.01% | 1.092 | 1.178 | 5.38% | 50.65% |
| 0.1 | 0.932 | 1.886 | 1.898 | 88.71% | 45.01% | 1.129 | 1.179 | 5.28% | 51.61% |
| 0.2 | 0.960 | 1.876 | 1.931 | 89.03% | 44.94% | 1.099 | 1.194 | 5.34% | 51.12% |
| 0.5 | 0.951 | 1.866 | 1.915 | 89.26% | 45.01% | 1.070 | 1.183 | 5.26% | 51.66% |
| 1 | 0.941 | 1.900 | 1.891 | 88.52% | 45.04% | 1.119 | 1.174 | 5.44% | 50.38% |