Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive, a failure-guided post-training framework that learns from self-generated failures through targeted supervision and safety-aware reinforcement learning. Reliable Diagnosis derives structured, verifiable feedback on collisions and drivable-area violations directly from simulator states. Minimum-Correction Target Retrieval searches a clustered human trajectory bank for nearby corrections that satisfy hard-safety constraints in the current scene, prioritizing preservation of the failed prediction's motion pattern. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the retrieved correction as a training-only auxiliary task. We then apply GRPO with a Safety-Layered Reward that strictly prioritizes hard-safe trajectories, retains continuous safety feedback for both unsafe and hard-safe trajectories, and rewards driving progress only after hard safety is satisfied. At inference, the policy directly predicts trajectories from the driving context without an explicit diagnosis or repair stage. On NAVSIM v1, RefineDrive improves the 4B base SFT policy from 87.7 to 91.7 PDMS. Using the same checkpoint without additional training, RefineDrive achieves 89.4 EPDMS on the original NAVTEST scenes evaluated with NAVSIM v2 extended metrics. Controlled ablations support the benefits of structured diagnosis supervision, retrieved corrections, and safety-layered optimization for direct planning.
Figures & tables
Figure 1: Motivation and overview of RefineDrive. Existing failure-learning methods can suffer from unreliable VLM-based diagnoses, over-corrective human ground-truth supervision, and coarse aggregated rewards. We address these limitations with Reliable Diagnosis, which grounds structured feedback in NAVSIM PDM simulation; Minimum-Correction Target Retrieval, which searches a clustered human trajectory bank for safe corrections close to failed predictions; and Safety-Layered Reward, which provides hierarchical signals over hard safety, continuous safety, and driving progress.
Figure 2: Training pipeline of RefineDrive. Offline rollouts and simulator-based safety checks expose model-specific failures. For each failure, we construct a structured diagnosis and retrieve a nearby safe correction from a human trajectory bank. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the correction. Finally, GRPO refines the policy with the Safety-Layered Reward, providing fine-grained feedback beyond aggregated driving scores. Diagnosis and correction are training-only tasks; inference directly maps the driving context to the final trajectory.
Method
Params.
NC
DAC
TTC
C
EP
PDMS
End-to-End Methods
UniAD [ 19 ]
–
97.8
91.9
92.9
100
78.8
83.4
TransFuser [ 20 ]
–
97.7
92.8
92.8
100
79.2
84.0
DiffusionDrive [ 21 ]
–
98.2
96.2
94.7
100
82.2
88.1
Driving VLA Methods
Base SFT (Qwen3.5-4B) [ 22 ]
4B
98.6
95.8
95.4
100
81.1
87.7
Table 1: Comparison with state-of-the-art methods on NAVSIM v1. All metrics are higher-is-better.
Method
NC
DAC
DDC
TLC
EP
TTC
LK
HC
EC
EPDMS
ReCogDrive [ 9 ]
98.3
95.2
99.5
99.8
87.1
97.5
96.6
98.3
86.5
83.6
DiffusionDriveV2* [ 27 ]
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
85.5
DriveVLA-W0* [ 23 ]
98.5
99.1
98.0
99.7
86.4
98.1
93.2
97.9
58.9
86.1
DriveSuprim (ViT-L) [ 28 ]
98.4
98.6
99.6
99.8
90.5
97.8
97.0
98.3
78.6
87.1
ExploreVLA* [ 29 ]
98.8
96.2
99.6
99.8
87.1
98.2
97.8
98.3
86.8
88.8
DriveFine† [ 26 ]
98.7
97.3
99.5
99.8
88.7
97.8
97.7
98.4
83.8
89.7
Table 2: Comparison on the original NAVTEST scenes using NAVSIM v2 extended metrics. Our result uses the same checkpoint as the NAVSIM v1 evaluation, without additional training. All metrics are higher-is-better. * indicates the exact NAVSIM v2 evaluator revision is not specified in the paper. † indicates results evaluated on the bug-fixed version of NAVSIM.
Correction Strategy
NC
DAC
EP
TTC
C
PDMS
Base SFT
98.6
95.8
81.1
95.4
100
87.7
Human GT
98.6
96.1
81.6
95.7
100
88.2
Minimal w/o Diagnosis
98.5
96.2
81.7
95.7
100
88.3
Minimal
98.7
96.4
82.0
95.9
100
88.6
Table 3: Ablation of failure-aware Correction SFT. Human GT and Minimal use the same failure cases and identical structured diagnosis supervision. Minimal w/o Diagnosis retains the same corrective trajectory targets as Minimal but removes diagnosis supervision.
RL Objective
NC
DAC
EP
TTC
C
PDMS
No RL
98.7
96.4
82.0
95.9
100
88.6
PDMS Reward
98.2
97.8
87.2
94.9
100
91.1
Safety-Layered Reward
98.7
98.2
87.2
95.9
100
91.7
Table 4: Ablation of the RL objective from the same minimal-correction SFT checkpoint.
Training Variant
NC
DAC
EP
TTC
C
PDMS
w/o Correction SFT
97.8
97.7
88.8
94.1
100
91.2
Minimal-Correction SFT w/o Diagnosis
98.2
97.8
87.5
95.3
99.9
91.3
Human-GT Correction SFT
98.1
97.7
89
94.7
99.9
91.6
Full Correction-SFT
98.7
98.2
87.2
95.9
100
91.7
Table 5: Effect of Correction SFT under the same Safety-Layered GRPO objective. Full Correction SFT achieves the highest NC, DAC, and TTC among the compared variants, with lower EP.
Figure 3: Qualitative comparison of direct trajectory predictions on NAVTEST: (a) off-road cases and (b) collision cases. Red, blue, and green denote the base-policy prediction, RefineDrive prediction, and human trajectory, respectively. Both policies predict independently from the same driving context; RefineDrive is not conditioned on the base-policy prediction.
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and planning. However, existing approaches lack mechanisms to exploit past failures or adapt to distribution shifts, causing the model to persistently underperform on similar scenarios where it has previously failed. In this paper, we propose DriveVLA-M0, a retrieval-augmented VLA with failure-aware latent memory. We construct a latent memory pool that stores failure cases along with their structure scene representations and expert trajectory labels, and design a dedicated Retrieve Model that decouples static road structure and dynamic agent interactions to enable structurally grounded retrieval. At inference time, retrieved cases are injected into the model via a lightweight decoupled LoRA-based test-time training (TTT) mechanism, allowing targeted and scenario-specific correction without modifying the backbone. Extensive experiments on NAVSIMv1 and NAVSIMv2 benchmark demonstrate that our approach consistently outperforms prior methods, achieving 94.1 PDMS on Navtest and 47.0 EPDMS on Navhard with only 26.44 ms TTT backward latency overhead. Furthermore, we show that DriveVLA-M0 scales effectively with additional memory, enabling training-free performance gains through memory expansion. The code is available at https://github.com/ZebinX/DriveVLA-M0.
Zebin Xing, Yupeng Zheng, Qiang Chen +10
Institute of Automation, Chinese Academy of Sciences · Chongqing Chang’an Technology Co., Ltd.
Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectories sampled from the current policy. Group relative policy optimization (GRPO) learns from reward differences within each rollout group. When all sampled trajectories are poor, this relative signal can rank failures without identifying behavior outside the failed region. We introduce FIRE-VLA, a failure-informed self-evolution framework that converts such unresolved failures into privileged supervision for the next policy. Low-reward, low-diversity groups trigger self-distillation from a frozen round-start copy of the same model. Teacher and student have the same parameter scale, but only the teacher observes the hidden future trajectory. Supervision follows the student's generated prefix and is restricted to answer tokens, while GRPO remains active for every group. The updated policy supplies the teacher for the next round, allowing the routed failure distribution to change with the policy without requiring a larger external teacher. Starting from the same Qwen2.5-VL-3B SFT checkpoint, the comparison matches student rollout and policy-update counts. On 6,019 examples from 150 held-out nuScenes scenes, FIRE-VLA retains comparable single-sample planning, reduces G=4 mean L2 from 1.848 to 1.500 m, and lowers evaluation-persistent failure prevalence from 13.03% to 11.20%. The reduction in mean error arises mainly from rare severe rollouts rather than uniform improvement across ordinary trajectories.
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes linguistic reasoning rather than action-grounded planning. As a result, the learned representations capture semantic knowledge but lack spatial dependencies crucial for reliable trajectory prediction. We propose DriveTeach-VLA, a framework that explicitly teaches VLAs what to see and where to look. Driving-aware Vision Distillation (DVD) injects driving-specific perceptual priors into the vision encoder, while 2D Trajectory-Guided Prompts (2D-TGP) provide spatial conditioning aligned with feasible driving trajectories. Together, they form a vision-guided learning pipeline: what to see (DVD pretraining) - where to look (TGP-guided SFT) - how to act (TGP-guided GRPO). DriveTeach-VLA achieves the state-of-the-art performance on NAVSIM and nuScenes. Our code is available at: https://github.com/ShivaTeam/DriveTeach-VLA.
Yuguang Yang, Canyu Chen, Zhewen Tan +10
School of Electronic Information Engineering, Beihang University · Institute for AI Industry Research (AIR), Tsinghua University · National College for Excellent Engineers, Beihang University +5