Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
Figures & tables
Figure 1: An illustration of AD-Memo. AD-Memo generates language-based memory periodically, which is beneficial for a diverse set of tasks illustrated in panel (a), (b) and (c), including driving, scene understanding question-answering tasks, and producing portable memory for other models.
Figure 2: An illustration of supervised training data of AD-Memo; the three tasks listed in panel (a), (b) and (c) are jointly trained in the SFT stage (all-way stop dataset only has (a) and (b)). Driving and memory-writing data are one entry per step, while question-answering data is one entry per clip.
Figure 3: An illustration of semi-closed loop RL, where we keep K parallel memories; the memory and driving tokens require heterogeneous handling, with the former being trajectory-wise and the latter being stepwise.
minADE ↓
Avg. ADE ↓
ML. ADE ↓
Stop SR ↑
Go SR ↑
Δpos↓
Δdur↓
Roll ↓
MCQ Acc. ↑
Base Model
1.386
2.351
2.393
82.06%
33.74%
1.725
1.871
11.79%
0%
Alpamayo 2 Super 34B
0.993
2.293
N/A
78.75%
34.30%
2.182
1.488
15.94%
33.19%
No mem. (SFT only)
1.102
2.185
2.318
87.04%
38.59%
1.358
1.564
6.62%
33.77%
CoT mem. (SFT only)
1.096
2.171
2.296
86.89%
39.37%
1.377
1.531
6.69%
45.41%
AD-Memo (SFT only)
1.049
2.121
2.256
87.48%
40.66%
1.257
1.469
6.05%
45.99%
CoT mem. (SFT+ref. mem)
1.098
2.169
2.286
86.95%
39.45%
1.367
1.521
6.61%
45.24%
Table 1: The main result for all-way stop dataset, with the best results bolded. “ref. mem” stands for using “reference memory” as input; they are generated in the same way as SFT ground truth labels. ↓ means the lower the better, and vice versa. The result shows that AD-Memo works the best on both driving and question-answering tasks, and even far outperforms the flagship model, Alpamayo2 Super 34B (which has action expert and does not output logits, thus has no most likely ADE).
Figure 4: An example of how memory aids driving. The ego and the car on the left (circled in red) are both stopping. Without memory, the agent does not know whether it should move first.
Figure 5: The trajectory of different methods with the input from Fig. 4 . The result shows that our method can better recognize that ego should stop, which yields lower ADE results.
minADE ↓
Avg. ADE ↓
ML. ADE ↓
Stop SR ↑
Go SR ↑
Δpos↓
Δdur↓
Roll ↓
MCQ Acc. ↑
Da Capo (Ours)
0.951
1.866
1.915
89.26%
45.01%
1.070
1.183
5.26%
51.66%
Capo (no std)
0.942
1.952
2.047
89.22%
44.27%
1.130
1.277
5.01%
50.91%
Capo + const. scaling
0.943
1.911
1.979
89.21%
44.74%
1.122
1.227
5.03%
51.57%
GRPO (traj-level reward)
1.007
2.113
2.194
87.11%
40.94%
1.312
1.392
6.28%
48.16%
AD-Memo (SFT only)
1.049
2.121
2.256
87.48%
40.66%
1.257
1.469
6.05%
45.99%
Table 2: The ablation on Da Capo , which proves that a better credit assignment significantly improves final performance over trajectory-level GRPO. The standard deviation in advantage also helps by introducing a natural scaling factor that balances rewards of different magnitudes.
minADE ↓
Avg. ADE ↓
ML. ADE ↓
minFDE ↓
Avg. FDE ↓
ML. FDE ↓
Corner ↓
MCQ Acc. ↑
Base Model
1.032
1.981
2.049
2.853
5.937
6.194
1.001
0
Alpamayo 2 Super 34B
0.954
2.081
N/A
2.556
6.141
N/A
0.902
25.71%
No mem. (SFT only)
1.007
2.033
2.075
2.716
6.108
6.279
0.968
58.02%
CoT mem. (SFT only)
1.031
2.059
2.105
2.787
6.192
6.375
0.993
57.51%
AD-Memo (SFT only)
1.005
2.011
2.049
2.705
6.042
6.196
0.967
65.63%
CoT mem. (SFT+ref. mem)
1.034
2.064
2.105
2.795
6.207
6.373
0.995
57.36%
Table 3: The main result for the general driving dataset, where the best results are bolded; the result shows that our method performs the best. ADE, FDE and corner distance are all the lower the better; the MCQ accuracy is the higher the better.
Figure 6: An image paired with the question “What is the color of traffic lights?” Repeated statements in the unmerged memory that “the ego vehicle is now stopped” mislead GPT-5.6 Luna into incorrectly inferring that the traffic light was red.
Our memory + last frame
Unmerged memory + last frame
Only last frame
Full video
WaymoQA
71.65%
68.3%
66.63%
75%
LingoQA
67%
60.4%
62.6%
70%
Table 4: The accuracy of GPT-5.6 Luna on LingoQA and WaymoQA, which shows that our memory can enhance the scene understanding of other models in a plug-and-play manner.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Supervised Fine-Tuning
global batch size
128
learning rates
1×10−5 (vision encoder); 1×10−4 (language model and LM head)
optimizer
fused AdamW
Adam betas / epsilon
(0.9,0.95) / 1×10−8
weight decay
0.1
lr scheduler
cosine decay; 100-step linear warm-up; minimum lr factor 0.01
Appendix
Table 5: List of hyperparameters for our algorithm.
Figure 7: An illustration of the two key statistics of our dataset. Panel (a) is the number of keyframes, which is the length of our clip; panel (b) is the co-wait time between ego and other car, which is a surrogate for the memory horizon required.
Figure 8: An example of the visual input for all-way stop dataset at the VQA frame.
Figure 9: An example of a decision graph.
Figure 10: The distributions of the number of keyframes in the general driving dataset. Most clips are around 21 frames, which translates to 8 seconds of video.
Figure 11: An example of the visual input for the general driving dataset at the VQA frame.
Figure 12: An example where memory benefits general driving; a pedestrian in the red box is occluded behind the crosswalk sign, which is hard to discover; the intention of the pedestrian is even harder to tell.
Figure 13: The previous frames of the same clip as that in Fig. 12 . With these frames, it is clearly shown that the pedestrian is walking towards the road and the ego car should yield.
Figure 14: The trajectories predicted by different methods in the example of Fig. 12 and Fig. 13 . It is clearly shown that our method outperforms baselines.
minADE ↓
Avg. ADE ↓
ML. ADE ↓
Stop SR ↑
Go SR ↑
Δpos↓
Δdur↓
Roll ↓
MCQ Acc. ↑
Alpamayo 2 Super 34B
0.993
2.293
N/A
78.75%
34.30%
2.182
1.488
15.94%
33.19%
Base model (2B)
1.386
2.351
2.393
82.06%
33.74%
1.725
1.871
11.79%
0%
No memory (2B SFT)
1.102
2.185
2.318
87.04%
38.59%
1.358
1.564
6.62%
33.77%
CoT-as-memory (2B SFT)
1.096
2.171
2.296
86.89%
39.37%
1.377
1.531
6.69%
45.41%
Ours (2B SFT)
1.049
2.121
2.256
87.48%
40.66%
1.257
1.469
6.05%
45.99%
Base model (8B)
1.315
2.254
2.332
83.72%
35.65%
1.627
1.776
10.90%
0%
Appendix
Table 6: Results of 8B models on All-Way-Stop v4. The result clearly shows that our 8B model works the best.
Easy
Hard (#1, past-anchor)
Hard (#2, multi-hop)
Avg. test set
Random
16.67%
16.67%
16.67%
16.67%
Alpamayo 2 Super 34B
48.45%
25.77%
25.65%
25.71%
No memory
98.14%
54.93%
61.10%
58.02%
CoT-as-Memory
98.02%
51.46%
63.56%
57.51%
Ours
97.92%
60.92%
70.34%
65.63%
CoT-as-Memory (GT mem)
97.99%
51.25%
63.48%
57.36%
Appendix
Table 7: The MCQ accuracy on our “easy” VQA questions (i.e. the one in the training set) and “hard” VQA questions (i.e. the one in the test set) for the general driving dataset. GPT-5.6 Luna with full video serves as a sanity check and verifier for the correctness of our generated VQA questions. Avg. test set is the average of hard #1 and hard #2; the example for each type is listed in Sec. C.2 . It is clearly shown that test on the same distribution of VQA questions as training causes overfitting. Thus, we deliberately use different questions for the training and test set.
Figure 15: The distribution of different scenes in the test set of our general driving dataset.
ADE gain
FDE gain
Shares
(min/Avg./ML)
(min/Avg./ML)
(%)
Vulnerable Road Users
-0.011 /+0.010/ +0.018
-0.011 / +0.054 / +0.072
41.46%
Traffic Control Compliance
+0.013 / +0.023 / +0.030
+0.089 / +0.101 / +0.105
20.57%
Lead Vehicle Following
+0.003/ +0.017 / +0.012
+0.021 / +0.058 / +0.031
8.98%
Turning Maneuver
-0.002/ +0.024 / +0.031
-0.019 / +0.061 / +0.092
7.81%
Stop for Vehicles
+0.008/ +0.011 / +0.017
+0.057 / +0.072 / +0.067
7.79%
Appendix
Table 8: Per-scenario open-loop trajectory error gains of our model over the no-memory model (after RL); as it is the gain , the higher the better, i.e., positive values indicate improvement , and vice versa. The categories are sorted in descending order of their shares. All values are in meters. It is shown that AD-Memo outperforms the no-memory model on the more common scenes, especially traffic control compliance, lead vehicle following and stop for vehicles.
Figure 16: An example of how our pipeline based on decision graph identify all objects related to driving. All objects that affect driving decision are grounded in panel (b) from the input frame shown in panel (a); these objects end up as nodes in the corresponding decision graph in panel (c).
Model
Decode success rate
All-way stop dataset
Base model
100.00%
Alpamayo 2 Super 34B
100.00%
No memory (SFT)
99.44%
CoT-as-Memory (SFT)
99.41%
Ours (SFT)
99.51%
Appendix
Table 9: The decode success rate for our model and baselines, which are all reasonably close to 100%.
Figure 17: The training dynamics of our semi-closed loop RL on the all-way stop dataset. We report the original curve and Exponential Moving Average (EMA).
Figure 18: The training dynamics of our semi-closed loop RL on the general driving dataset. We report the original curve and Exponential Moving Average (EMA). The VQA accuracy is close to 1 as the problems are relatively easy to overfit, and thus we test different problems in the test set; see Sec. D.3 for details.
Figure 19: The change of performance with respect to memory horizon during inference time on the all-way stop dataset.
Figure 20: The change of performance with respect to memory horizon during inference time on the general driving dataset.
λVQA
minADE ↓
Avg. ADE ↓
ML. ADE ↓
Stop SR ↑
Go SR ↑
Δpos↓
Δdur↓
Roll ↓
MCQ Acc. ↑
0.02
0.945
1.886
1.927
88.73%
44.80%
1.111
1.184
5.54%
50.43%
0.05
0.943
1.877
1.913
88.67%
45.01%
1.092
1.178
5.38%
50.65%
0.1
0.932
1.886
1.898
88.71%
45.01%
1.129
1.179
5.28%
51.61%
0.2
0.960
1.876
1.931
89.03%
44.94%
1.099
1.194
5.34%
51.12%
0.5
0.951
1.866
1.915
89.26%
45.01%
1.070
1.183
5.26%
51.66%
1
0.941
1.900
1.891
88.52%
45.04%
1.119
1.174
5.44%
50.38%
Appendix
Table 10: The result of our method on the all-way stop dataset with different λVQA . The performance is generally consistent, and no hyperparameter choice prevail on all metrics.
Figure 21: The performance change with the training length of SFT on the validation set. It is clearly shown that, after 2 epochs, not only does the trajectory move further from ground truth, but the format following ability of the model is also degrading due to exposure bias ( Bengio et al., 2015 ) , i.e., more invalid rollouts appear in sampling.