cs.AISep 30, 2026

AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation

Authors: Minrui Liu, Jingke Wang, Yuehao Huang, Hao Su, Jiajun Lv, Yukai Ma, Yong Liu

Organizations: Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China.

Abstract

Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.

Figures & tables

Explore similar work

CardsList
  1. ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

    Jul 14, 2026Jiahang Wang, Yirong Yang, Yanqing Zhu +4Vision-Language NavigationNavigation

  2. Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

    Sep 29, 2026Zhimin Wang, Meiyuan Zhu, Duo Wu +8Vision-Language NavigationNavigation

  3. Joint On-and-Off Policy Learning for Vision-and-Language Navigation

    Jul 15, 2026Qingrong He, Lin Zhao, Kevin Zheng +1Vision-Language NavigationEmbodied Agents