AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation
Organizations: Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China.
Abstract
Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.
Figures & tables
| 22 25 23 26 24 27-28 Method | Observation | R2R-CE Val-Unseen | RxR-CE Val-Unseen | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pano. | Odo. | Depth | S.RGB | NE | OS | SR | SPL | NE | SR | SPL | nDTW | |
| HPN+DN ∗ [ 26 ] | ✓ | ✓ | ✓ | 6.31 | 40.0 | 36.0 | 34.0 | – | – | – | – | |
| CMA ∗ [ 27 ] | ✓ | ✓ | ✓ | 6.20 | 52.0 | 41.0 | 36.0 | 8.76 | 26.5 | 22.1 | 47.0 | |
| GridMM ∗ [ 28 ] | ✓ | ✓ | ✓ | 5.11 | 61.0 | 49.0 | 41.0 | – | – | – | – | |
| ETPNav ∗ [ 29 ] | ✓ | ✓ | ✓ | 4.71 | 65.0 | 57.0 | 49.0 | 5.64 | 54.7 | 44.8 | 61.9 | |
| ScaleVLN ∗ [ 30 ] | ✓ | ✓ | ✓ | 4.80 | – | 55.0 | 51.0 | – | – | – | – | |
| 6 Configuration | NE | SR | OS | nDTW |
|---|---|---|---|---|
| System 2 | 5.77 | 52.88 | 65.30 | 64.16 |
| System 2 + SFT | 5.69 | 53.41 | 65.30 | 64.20 |
| System 2 + DPO | 5.76 | 55.03 | 65.56 | 63.21 |
| System 2 + Monitor | 4.53 | 57.78 | 68.71 | 57.25 |
| System 2 + DPO + Monitor | 4.67 | 60.21 | 70.29 | 57.13 |