Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from 31.03% to 11.33%. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield 15.27% SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting Weak-trained translator raises Weak-input SR to 37.93% and transfers zero-shot to real human instructions (11.33%→32.51%); it also improves held-out OpenFly (4.95%→20.79%) and yields recovery on CityNav and AirVLN.
Figures & tables
Figure 1: Instruction gap and recoverable activation on a frozen OpenFly-Agent. (a) Median instruction length collapses from Orig to Weak to Human. (b) On the 203-episode evaluation set (Selected-Seen), SR drops from 31.03% (Orig) to 15.27% (Weak) and 11.33% (Human). TGIT(W) recovers to 37.93% on Weak; the same adapter transfers zero-shot as TGIT(H) to 32.51% on Human.
Figure 2: The overview of TGIT. (a) Hard-case pool : keep episodes with S(π,Im)=1 and S(π,Iw)=0 so that capability exists but the weak interface fails. (b) Candidate generation : VLM translator fθ takes (Iw,O1:M) and emits N=3 executable rewrites; only LoRA A / B on the LM decoder is trainable. (c) Reward + multi-round LoRA–DPO : frozen rollouts score each yi by R=0.5SR+0.3SPL+0.2e−0.1dnorm (SR: success rate; SPL: path efficiency; dnorm : normalized final distance; blue flags: targets), form preferences y+≻y− , update A / B , then chain the adapter and resample hard cases for the next round. (d) Training–deployment contract : evaluation collapses to one forward pass—no reward, preference, or policy update (Eq. 2 ).
Method
OpenFly-Agent ( Gao et al. , 2025 )
Selected-Seen
Unselected-Seen
Unseen
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
NE ↓
Orig
31.03
31.66
116.56
20.05
16.82
77.42
11.39
29.82
172.77
Weak
15.27
37.36
82.48
14.11
13.47
202.27
4.95
8.71
90.25
Human
11.33
89.43
108.92
–
–
–
–
–
–
TGIT(H)
32.51
92.75
95.31
–
–
–
–
–
–
Table 1: Main navigation results on three frozen UAV agents. We report Success Rate (SR, %), success-conditional path efficiency (SPL, %), and Navigation Error (NE, meters). SPL is averaged only over successful episodes. TGIT(H) and TGIT(W) denote Human and Weak evaluation inputs, respectively. Human and TGIT(H) are evaluated only on OpenFly Selected-Seen (203 episodes) due to the high cost of human annotation; other Human-input cells are “–”. TGIT(H) reuses the Weak-trained TGIT(W) adapter zero-shot, without Human-specific training. TGIT(W) deltas are against Orig (negative NE is better).
Variant
SR (%) ↑ / SPL (%) ↑ / NE (m) ↓
Round 1
Round 2
Round 3
Round 4
Round 5 (Final)
Random, Full
15.76 / 27.22 / 90.96
30.54 / 17.14 / 119.48
18.72 / 28.70 / 92.97
21.18 / 18.92 / 95.25
18.72 / 41.83 / 83.85
Random, Pool
22.17 / 40.29 / 96.05
19.21 / 24.26 / 108.89
22.17 / 30.34 / 93.33
20.69 / 54.93 / 85.83
17.73 / 36.48 / 87.72
Selected, Full
31.03 / 15.97 / 119.20
33.50 / 21.68 / 110.30
28.08 / 23.35 / 101.37
31.03 / 31.79 / 103.51
29.06 / 24.27 / 105.22
Selected, Pool
21.18 / 16.13 / 133.56
25.12 / 48.11 / 105.44
29.56 / 33.07 / 83.43
31.53 / 37.41 / 76.58
37.93 / 24.85 / 72.08
Table 2: Multi-round ablation on data pool construction (OpenFly-Agent). Results are SR (%) / SPL (%) / NE (m) . Selected : original succeeds and weakened fails. Pool : maintain 100 candidates and sample 20 per round.
Figure 3: Hard cases, failure stage, and visual context. (a) Pool curves (Table 2 ); Selected-Pool Round 5 matches Table 1 . (b) Failure stages with success fractions matched to Table 1 ; TGIT early/late proportions use the nearest Selected-Pool diagnostic after renormalizing success to 77/203 . (c) f1 ∗ /f4 ∗ diagnostics; “full” is TGIT in Table 1 .
Preference ranking
Valid
SR ↑
SPL ↑
NE ↓
Success only
203/203
35.47
93.13
91.48
Success → SPL → NE
202/203
36.14
95.00
84.60
SPL → NE
201/203
32.34
91.24
97.01
Table 3: Reward Ranking ( N=100 , Round 5). Alternatives to the main Weighted form on a fixed candidate set. Main TGIT SR remains 37.93% (Table 1 ).
Translator
Valid
SR
SPL
NE
Qwen2.5-VL-7B
203/203
36.45
95.17
90.34
Qwen3-VL-4B
203/203
29.06
90.96
128.33
Qwen3-VL-2B
156/203
31.53
93.67
82.33 †
Qwen3.5-0.8B
193/203
35.47
93.01
111.05 †
Table 4: Translator backbone/scale comparison ( N=100 , M=4 ; diagnostic). SR uses fixed denominator 203 ; invalid outputs count as failures. † NE is valid-only. TGIT in Table 1 remains the primary result.
Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments. While recent advances have achieved progress through large-scale memory graphs and lookahead path planning, they remain limited by shallow instruction understanding and high computational cost. In particular, existing methods rely primarily on landmark descriptions, overlooking directional cues "a key source of spatial context in human navigation". In this work, we propose LookasideVLN, a new paradigm that exploits directional cues in natural language to achieve both more accurate spatial reasoning and greater computational efficiency. LookasideVLN comprises three core components: (1) an Egocentric Lookaside Graph (ELG) that dynamically encodes instruction-relevant landmarks and their directional relationships, (2) a Spatial Landmark Knowledge Base (SLKB) that provides lightweight memory retrieval from prior navigation experiences, and (3) a Lookaside MLLM Navigation Agent that aligns multimodal information from user instructions, visual observations, and landmark-direction information from ELG for path planning. Extensive experiments show that LookasideVLN significantly outperforms the state-of-the-art CityNavAgent, even with a single-level lookahead, demonstrating that leveraging directional cues is a powerful yet efficient strategy for Aerial VLN.
Yuwei Ning, Ganlong Zhao, Yipeng Qin +4
Sun Yat-sen University · Peng Cheng Laboratory · The Chinese University of Hong Kong +4
Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale vision-language backbones and dense visual histories, imposing substantial computation and memory costs that hinder onboard deployment. We propose LightVLN, a lightweight history-aware aerial VLN framework that combines a compact 0.5B language backbone with compact representations of both historical and current observations. LightVLN compresses each historical frame into a single token using visual features already computed by the policy. It further introduces history- and instruction-conditioned local aggregation to reduce the current observation from 256 to 32 visual tokens while preserving navigation-relevant spatial information. With up to 16 historical frames, the policy uses at most 48 observation-derived tokens. On the public OpenFly dataset, LightVLN achieves 50.93% Test-Seen and 36.14% Test-Unseen success rates (SR), outperforming the evaluated 7B language-backbone baselines on most reported metrics. It also achieves 25.83% SR on AerialVLN-S Val-Seen. In a reconstructed unseen campus, we deploy LightVLN on a DJI M350 RTK with an external Jetson Orin NX 16 GB for closed-loop onboard-compute real-to-sim hardware-in-the-loop (HIL) evaluation, achieving 14.61 Hz model inference and 11.13 Hz end-to-end decision updates. These results demonstrate the effectiveness and efficiency of LightVLN for aerial navigation.
Yiming Zhao, Tianshun Li, Jingle He +2
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous multi-step instructions over long horizons. Existing zero-shot methods remain limited, as they often rely on large base models, generic prompts, and loosely coordinated modules. In this work, we propose FineCog-Nav, a top-down framework inspired by human cognition that organizes navigation into fine-grained modules for language processing, perception, attention, memory, imagination, reasoning, and decision-making. Each module is driven by a moderate-sized foundation model with role-specific prompts and structured input-output protocols, enabling effective collaboration and improved interpretability. To support fine-grained evaluation, we construct AerialVLN-Fine, a curated benchmark of 300 trajectories derived from AerialVLN, with sentence-level instruction-trajectory alignment and refined instructions containing explicit visual endpoints and landmark references. Experiments show that FineCog-Nav consistently outperforms zero-shot baselines in instruction adherence, long-horizon planning, and generalization to unseen environments. These results suggest the effectiveness of fine-grained cognitive modularization for zero-shot aerial navigation. Project page: https://smartdianlab.github.io/projects-FineCogNav.
Dian Shao, Zhengzheng Xu, Peiyang Wang +4
Northwestern Polytechnical University · Nanjing University