cs.ROOct 7, 2026

AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation

Authors: Zhenxuan Zeng, Qingle Wu, Wei Suo, Maojia Wu, Bairong Zhang, Hangzheng Yu, Peng Wang

Organizations: School of Computer Science, Northwestern Polytechnical University, China.

Abstract

Goal-oriented Vision-and-Language Navigation (VLN) requires agents to locate and reach targets described in natural language without prescribed routes. Air--ground collaboration is valuable for tasks requiring both wide-area search and fine-grained localization. However, systematic study of goal-oriented air--ground collaborative VLN remains limited by the lack of large-scale, diverse benchmarks and two core challenges: 1) substantial differences between aerial and ground views, together with useful observations becoming unavailable as navigation proceeds, make it difficult to maintain spatially consistent context across platforms and over time; and 2) asymmetric spatial observability makes ground perception locally detailed but spatially limited and aerial perception broad but locally coarse, limiting the reliability of single-platform planning. To address these limitations, we introduce AirGroundVLN, a benchmark containing 10,281 navigation episodes and 955 target instances across 19 Unreal Engine environments, with seen/unseen splits and an aerial-visibility protocol for systematic evaluation. Alongside the benchmark, we propose AG-CoNAV, a trainable reference framework comprising two key components: Spatiotemporally Anchored Collaborative Memory (SACM) and Aerial-Guided Regional-to-Local Planning (AGRLP). SACM maintains and retrieves spatially consistent historical context across aerial and ground observations. Meanwhile, AGRLP combines regional aerial guidance with fine-grained ground navigation. Extensive experiments demonstrate the effectiveness of AG-CoNAV and establish AirGroundVLN as a comprehensive benchmark for future exploration.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps

    Sep 3, 2026Shuning Zhang, Liang Li, Yunheng Wang +3Aerial Vision-Language NavigationCooperative Air Combat

  2. CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning

    Aug 3, 2026Junru Song, Wenhao Zhang, Yang Yang +7Aerial Vision-Language NavigationMulti-Uav

  3. From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

    Jul 4, 2026Xiangyu Shi, Ruoxi Yang, Wei Tao +3Vision-Language NavigationContextual Grounding