Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from 31.03% to 11.33%. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield 15.27% SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting Weak-trained translator raises Weak-input SR to 37.93% and transfers zero-shot to real human instructions (11.33%→32.51%); it also improves held-out OpenFly (4.95%→20.79%) and yields recovery on CityNav and AirVLN.
Figures & tables
Figure 1: Instruction gap and recoverable activation on a frozen OpenFly-Agent. (a) Median instruction length collapses from Orig to Weak to Human. (b) On the 203-episode evaluation set (Selected-Seen), SR drops from 31.03% (Orig) to 15.27% (Weak) and 11.33% (Human). TGIT(W) recovers to 37.93% on Weak; the same adapter transfers zero-shot as TGIT(H) to 32.51% on Human.
Figure 2: The overview of TGIT. (a) Hard-case pool : keep episodes with S(π,Im)=1 and S(π,Iw)=0 so that capability exists but the weak interface fails. (b) Candidate generation : VLM translator fθ takes (Iw,O1:M) and emits N=3 executable rewrites; only LoRA A / B on the LM decoder is trainable. (c) Reward + multi-round LoRA–DPO : frozen rollouts score each yi by R=0.5SR+0.3SPL+0.2e−0.1dnorm (SR: success rate; SPL: path efficiency; dnorm : normalized final distance; blue flags: targets), form preferences y+≻y− , update A / B , then chain the adapter and resample hard cases for the next round. (d) Training–deployment contract : evaluation collapses to one forward pass—no reward, preference, or policy update (Eq. 2 ).
Method
OpenFly-Agent ( Gao et al. , 2025 )
Selected-Seen
Unselected-Seen
Unseen
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
NE ↓
Orig
31.03
31.66
116.56
20.05
16.82
77.42
11.39
29.82
172.77
Weak
15.27
37.36
82.48
14.11
13.47
202.27
4.95
8.71
90.25
Human
11.33
89.43
108.92
–
–
–
–
–
–
TGIT(H)
32.51
92.75
95.31
–
–
–
–
–
–
Table 1: Main navigation results on three frozen UAV agents. We report Success Rate (SR, %), success-conditional path efficiency (SPL, %), and Navigation Error (NE, meters). SPL is averaged only over successful episodes. TGIT(H) and TGIT(W) denote Human and Weak evaluation inputs, respectively. Human and TGIT(H) are evaluated only on OpenFly Selected-Seen (203 episodes) due to the high cost of human annotation; other Human-input cells are “–”. TGIT(H) reuses the Weak-trained TGIT(W) adapter zero-shot, without Human-specific training. TGIT(W) deltas are against Orig (negative NE is better).
Variant
SR (%) ↑ / SPL (%) ↑ / NE (m) ↓
Round 1
Round 2
Round 3
Round 4
Round 5 (Final)
Random, Full
15.76 / 27.22 / 90.96
30.54 / 17.14 / 119.48
18.72 / 28.70 / 92.97
21.18 / 18.92 / 95.25
18.72 / 41.83 / 83.85
Random, Pool
22.17 / 40.29 / 96.05
19.21 / 24.26 / 108.89
22.17 / 30.34 / 93.33
20.69 / 54.93 / 85.83
17.73 / 36.48 / 87.72
Selected, Full
31.03 / 15.97 / 119.20
33.50 / 21.68 / 110.30
28.08 / 23.35 / 101.37
31.03 / 31.79 / 103.51
29.06 / 24.27 / 105.22
Selected, Pool
21.18 / 16.13 / 133.56
25.12 / 48.11 / 105.44
29.56 / 33.07 / 83.43
31.53 / 37.41 / 76.58
37.93 / 24.85 / 72.08
Table 2: Multi-round ablation on data pool construction (OpenFly-Agent). Results are SR (%) / SPL (%) / NE (m) . Selected : original succeeds and weakened fails. Pool : maintain 100 candidates and sample 20 per round.
Figure 3: Hard cases, failure stage, and visual context. (a) Pool curves (Table 2 ); Selected-Pool Round 5 matches Table 1 . (b) Failure stages with success fractions matched to Table 1 ; TGIT early/late proportions use the nearest Selected-Pool diagnostic after renormalizing success to 77/203 . (c) f1 ∗ /f4 ∗ diagnostics; “full” is TGIT in Table 1 .
Preference ranking
Valid
SR ↑
SPL ↑
NE ↓
Success only
203/203
35.47
93.13
91.48
Success → SPL → NE
202/203
36.14
95.00
84.60
SPL → NE
201/203
32.34
91.24
97.01
Table 3: Reward Ranking ( N=100 , Round 5). Alternatives to the main Weighted form on a fixed candidate set. Main TGIT SR remains 37.93% (Table 1 ).
Translator
Valid
SR
SPL
NE
Qwen2.5-VL-7B
203/203
36.45
95.17
90.34
Qwen3-VL-4B
203/203
29.06
90.96
128.33
Qwen3-VL-2B
156/203
31.53
93.67
82.33 †
Qwen3.5-0.8B
193/203
35.47
93.01
111.05 †
Table 4: Translator backbone/scale comparison ( N=100 , M=4 ; diagnostic). SR uses fixed denominator 203 ; invalid outputs count as failures. † NE is valid-only. TGIT in Table 1 remains the primary result.