cs.AIOct 7, 2026

Speaking the Navigator's Language: Trajectory-Grounded Instruction Translation for Frozen Aerial VLN Agents

Authors: Xi Chen, Zhe Liu, Xiaogang Xu, Jiafei Xu, Chunyi Zhou, Yuan Su, Rui Zeng, Tianyu Du, +3 more

Organizations: Zhejiang University · Zhejiang Lab

Abstract

Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from 31.03%31.03\% to 11.33%11.33\%. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield 15.27%15.27\% SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting Weak-trained translator raises Weak-input SR to 37.93%37.93\% and transfers zero-shot to real human instructions (11.33%→32.51%11.33\%{\rightarrow}32.51\%); it also improves held-out OpenFly (4.95%→20.79%4.95\%{\rightarrow}20.79\%) and yields recovery on CityNav and AirVLN.

Figures & tables

Explore similar work

CardsList
  1. LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

    Apr 19, 2026Yuwei Ning, Ganlong Zhao, Yipeng Qin +4Vision-Language NavigationAerial VLN

  2. LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation

    Oct 4, 2026Yiming Zhao, Tianshun Li, Jingle He +2Memory-Augmented VLMsVision-Language Models

  3. FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

    Apr 17, 2026Dian Shao, Zhengzheng Xu, Peiyang Wang +4Zero-Shot LearningUAV Navigation