Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive, a failure-guided post-training framework that learns from self-generated failures through targeted supervision and safety-aware reinforcement learning. Reliable Diagnosis derives structured, verifiable feedback on collisions and drivable-area violations directly from simulator states. Minimum-Correction Target Retrieval searches a clustered human trajectory bank for nearby corrections that satisfy hard-safety constraints in the current scene, prioritizing preservation of the failed prediction's motion pattern. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the retrieved correction as a training-only auxiliary task. We then apply GRPO with a Safety-Layered Reward that strictly prioritizes hard-safe trajectories, retains continuous safety feedback for both unsafe and hard-safe trajectories, and rewards driving progress only after hard safety is satisfied. At inference, the policy directly predicts trajectories from the driving context without an explicit diagnosis or repair stage. On NAVSIM v1, RefineDrive improves the 4B base SFT policy from 87.7 to 91.7 PDMS. Using the same checkpoint without additional training, RefineDrive achieves 89.4 EPDMS on the original NAVTEST scenes evaluated with NAVSIM v2 extended metrics. Controlled ablations support the benefits of structured diagnosis supervision, retrieved corrections, and safety-layered optimization for direct planning.
Figures & tables
Figure 1: Motivation and overview of RefineDrive. Existing failure-learning methods can suffer from unreliable VLM-based diagnoses, over-corrective human ground-truth supervision, and coarse aggregated rewards. We address these limitations with Reliable Diagnosis, which grounds structured feedback in NAVSIM PDM simulation; Minimum-Correction Target Retrieval, which searches a clustered human trajectory bank for safe corrections close to failed predictions; and Safety-Layered Reward, which provides hierarchical signals over hard safety, continuous safety, and driving progress.
Figure 2: Training pipeline of RefineDrive. Offline rollouts and simulator-based safety checks expose model-specific failures. For each failure, we construct a structured diagnosis and retrieve a nearby safe correction from a human trajectory bank. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the correction. Finally, GRPO refines the policy with the Safety-Layered Reward, providing fine-grained feedback beyond aggregated driving scores. Diagnosis and correction are training-only tasks; inference directly maps the driving context to the final trajectory.
Method
Params.
NC
DAC
TTC
C
EP
PDMS
End-to-End Methods
UniAD [ 19 ]
–
97.8
91.9
92.9
100
78.8
83.4
TransFuser [ 20 ]
–
97.7
92.8
92.8
100
79.2
84.0
DiffusionDrive [ 21 ]
–
98.2
96.2
94.7
100
82.2
88.1
Driving VLA Methods
Base SFT (Qwen3.5-4B) [ 22 ]
4B
98.6
95.8
95.4
100
81.1
87.7
Table 1: Comparison with state-of-the-art methods on NAVSIM v1. All metrics are higher-is-better.
Method
NC
DAC
DDC
TLC
EP
TTC
LK
HC
EC
EPDMS
ReCogDrive [ 9 ]
98.3
95.2
99.5
99.8
87.1
97.5
96.6
98.3
86.5
83.6
DiffusionDriveV2* [ 27 ]
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
85.5
DriveVLA-W0* [ 23 ]
98.5
99.1
98.0
99.7
86.4
98.1
93.2
97.9
58.9
86.1
DriveSuprim (ViT-L) [ 28 ]
98.4
98.6
99.6
99.8
90.5
97.8
97.0
98.3
78.6
87.1
ExploreVLA* [ 29 ]
98.8
96.2
99.6
99.8
87.1
98.2
97.8
98.3
86.8
88.8
DriveFine† [ 26 ]
98.7
97.3
99.5
99.8
88.7
97.8
97.7
98.4
83.8
89.7
Table 2: Comparison on the original NAVTEST scenes using NAVSIM v2 extended metrics. Our result uses the same checkpoint as the NAVSIM v1 evaluation, without additional training. All metrics are higher-is-better. * indicates the exact NAVSIM v2 evaluator revision is not specified in the paper. † indicates results evaluated on the bug-fixed version of NAVSIM.
Correction Strategy
NC
DAC
EP
TTC
C
PDMS
Base SFT
98.6
95.8
81.1
95.4
100
87.7
Human GT
98.6
96.1
81.6
95.7
100
88.2
Minimal w/o Diagnosis
98.5
96.2
81.7
95.7
100
88.3
Minimal
98.7
96.4
82.0
95.9
100
88.6
Table 3: Ablation of failure-aware Correction SFT. Human GT and Minimal use the same failure cases and identical structured diagnosis supervision. Minimal w/o Diagnosis retains the same corrective trajectory targets as Minimal but removes diagnosis supervision.
RL Objective
NC
DAC
EP
TTC
C
PDMS
No RL
98.7
96.4
82.0
95.9
100
88.6
PDMS Reward
98.2
97.8
87.2
94.9
100
91.1
Safety-Layered Reward
98.7
98.2
87.2
95.9
100
91.7
Table 4: Ablation of the RL objective from the same minimal-correction SFT checkpoint.
Training Variant
NC
DAC
EP
TTC
C
PDMS
w/o Correction SFT
97.8
97.7
88.8
94.1
100
91.2
Minimal-Correction SFT w/o Diagnosis
98.2
97.8
87.5
95.3
99.9
91.3
Human-GT Correction SFT
98.1
97.7
89
94.7
99.9
91.6
Full Correction-SFT
98.7
98.2
87.2
95.9
100
91.7
Table 5: Effect of Correction SFT under the same Safety-Layered GRPO objective. Full Correction SFT achieves the highest NC, DAC, and TTC among the compared variants, with lower EP.
Figure 3: Qualitative comparison of direct trajectory predictions on NAVTEST: (a) off-road cases and (b) collision cases. Red, blue, and green denote the base-policy prediction, RefineDrive prediction, and human trajectory, respectively. Both policies predict independently from the same driving context; RefineDrive is not conditioned on the base-policy prediction.
School of Electronic Information Engineering, Beihang University · Institute for AI Industry Research (AIR), Tsinghua University · National College for Excellent Engineers, Beihang University +5