cs.CVSep 30, 2026

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

Authors: Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie

Organizations: Harbin Institute of Technology (Shenzhen) · Pengcheng Laboratory

Abstract

Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3% success and 1.51,m navigation error across eight challenging routes evaluated three times each.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

    Jul 14, 2026Jiahang Wang, Yirong Yang, Yanqing Zhu +4Vision-Language NavigationNavigation

  2. NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory

    Sep 28, 2026Kai Sheng, Liuyi Wang, Jinlong Li +3Vision-Language NavigationDiffusion-Based Vision-Language-Actions

  3. SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

    May 17, 2026Jingzhi Huang, Junkai Huang, Wenxuan Song +4Vision-Language NavigationMultimodal Large Language Models