cs.CVSep 28, 2026

RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving

Authors: Zhe Sun, Ziyi Luo, Yehao Lu, Lei Zhou, Xi Li

Organizations: College of Computer Science and Technology, Zhejiang University, Hangzhou, China · Yinwang Intelligent Technology Co., Ltd.

Abstract

Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive, a failure-guided post-training framework that learns from self-generated failures through targeted supervision and safety-aware reinforcement learning. Reliable Diagnosis derives structured, verifiable feedback on collisions and drivable-area violations directly from simulator states. Minimum-Correction Target Retrieval searches a clustered human trajectory bank for nearby corrections that satisfy hard-safety constraints in the current scene, prioritizing preservation of the failed prediction's motion pattern. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the retrieved correction as a training-only auxiliary task. We then apply GRPO with a Safety-Layered Reward that strictly prioritizes hard-safe trajectories, retains continuous safety feedback for both unsafe and hard-safe trajectories, and rewards driving progress only after hard safety is satisfied. At inference, the policy directly predicts trajectories from the driving context without an explicit diagnosis or repair stage. On NAVSIM v1, RefineDrive improves the 4B base SFT policy from 87.7 to 91.7 PDMS. Using the same checkpoint without additional training, RefineDrive achieves 89.4 EPDMS on the original NAVTEST scenes evaluated with NAVSIM v2 extended metrics. Controlled ablations support the benefits of structured diagnosis supervision, retrieved corrections, and safety-layered optimization for direct planning.

Figures & tables

Explore similar work

CardsList
  1. DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

    Aug 11, 2026Zebin Xing, Yupeng Zheng, Qiang Chen +10Autonomous DrivingDrives

  2. Teaching Vision-Language-Action Models What to See and Where to Look

    Jul 2, 2026Yuguang Yang, Canyu Chen, Zhewen Tan +10Diffusion-Based Vision-Language-ActionsAutonomous Driving