Organizations: State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China · Division of Natural and Applied Sciences, Duke Kunshan University, Suzhou, China · Peng Cheng Laboratory, ShenZhen, China
Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across different temporal horizons. To address this limitation, we propose ForeFly, a dual-horizon latent world action model that predicts both a proximal future for local continuity and an adaptive route-critical future for long-range guidance. Horizon-specific foresight queries are primed with recent and route-critical visual memories, providing history-aware context for future prediction. To exploit their distinct roles in action generation, we introduce Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance. Experiments on the TravelUAV and UAV-ON benchmarks show that ForeFly consistently outperforms strong baselines across seen and unseen settings, validating the effectiveness of dual-horizon foresight and FGAR learning. The code is available at: https://github.com/kunhuiW/ForeFly
Figures & tables
Figure 1: (a) Reactive policies predict actions without future modeling. (b) Single-horizon methods predict a single future latent, overlooking complementary navigation roles for local continuity and route-level anticipation. (c) ForeFly jointly predicts proximal and route-critical futures, while FGAR integrates them for locally smooth and strategically consistent navigation.
Figure 2: Overview of ForeFly. ForeFly primes dual-horizon foresight with complementary historical memories, predicts proximal and route-critical future latents, Foresight-Guided Action Refinement (FGAR) asymmetrically uses proximal foresight for local enhancement and route-critical foresight for feature-wise correction and route-level guidance.
Method
Seen
Unseen
SR ↑
OSR ↑
SPL ↑
SR ↑
OSR ↑
SPL ↑
AerialVLN Liu et al. (2023)
4.55
16.84
4.24
3.51
20.77
3.41
AOA-F Xiao et al. (2025)
7.87
17.98
4.21
6.45
16.09
3.98
OpenFly Gao et al. (2026)
12.29
26.20
6.83
11.82
25.71
5.64
APEX Zhang et al. (2026a)
12.36
19.10
9.13
16.13
22.58
13.03
Navid Zhang et al. (2024)
14.17
30.08
7.19
10.70
28.51
5.99
Table 4: Results on the UAV-ON dataset.
Method
Seen
Unseen
SR ↑
OSR ↑
SPL ↑
SR ↑
OSR ↑
SPL ↑
AerialVLN Liu et al. (2023)
4.55
16.84
4.24
3.51
20.77
3.41
AOA-F Xiao et al. (2025)
7.87
17.98
4.21
6.45
16.09
3.98
OpenFly Gao et al. (2026)
12.29
26.20
6.83
11.82
25.71
5.64
APEX Zhang et al. (2026a)
12.36
19.10
9.13
16.13
22.58
13.03
Navid Zhang et al. (2024)
14.17
30.08
7.19
10.70
28.51
5.99
Table 4: Results on the UAV-ON dataset.
Method
NE ↓
SR ↑
OSR ↑
SPL ↑
MLP Head
64.41
33.08
62.91
27.99
Direct Dual-Future*
62.03
34.00
62.50
28.73
FGAR w/o Btc
54.61
38.35
68.73
32.55
FGAR
48.74
41.75
70.25
34.75
Table 5: Ablation on Foresight-Guided Action Refinement (FGAR) Design.
Figure 3: Qualitative Results. (a) Bird’s-eye view of predicted trajectories. (b) 3D trajectory comparison showing improved route consistency of ForeFly. (c) Step-by-step navigation, where route-critical foresight helps ForeFly anticipate the upcoming turn, while the baseline deviates. The corresponding future visualization procedure is described in Sec. B.2 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Visualization of route-critical anchor construction. Representative UAV trajectories of different lengths are shown. Green and black stars mark the start and end states, red stars indicate key frames at turning points, and yellow diamonds denote distance anchors. The color scale represents the Euclidean distance from the starting position.
Figure 5: Results on TravelUAV Unseen Map and Unseen Object subsets. The top row compares ForeFly with existing methods on the Unseen Map subset, and the bottom row reports results on the Unseen Object subset.
Figure 6: Qualitative visualization of dual-horizon future prediction. Each row shows two sampled waypoints from the same navigation trajectory. For each waypoint, we visualize the current observation together with the ground-truth and predicted proximal and route-critical futures.
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination. To address these challenges, we propose DreamFly, a diffusion-based aerial VLN framework built on Dream-VLA. DreamFly introduces a causally aligned historical memory that augments the current visual representation using only observations preceding the current decision step, enabling temporal reasoning without future information leakage. We further formulate navigation as receding-horizon diffusion planning, where the policy predicts a K-step action chunk but executes only the first action before replanning. This plan-K, execute-one strategy uses future actions as auxiliary planning targets while preserving closed-loop visual feedback. Finally, LiteStop estimates the stop probability directly from action logits at the initial all-mask state, decoupling explicit termination from action generation. Experiments on the OpenFly benchmark demonstrate consistent improvements in seen and unseen environments. DreamFly achieves 32.04%/29.46% SR and 28.22%/23.54% SPL on the test-seen/test-unseen splits, respectively, outperforming all compared methods on both metrics while attaining the lowest navigation error. These results demonstrate the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial VLN.
Yan Deng, Fei Xu
School of Electronic Information Engineering, Xi’an Technological University, Xi’an, China · School of Computer Science and Engineering, Xi’an Technological University, Xi’an, China
Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet existing Vision-Language Navigation (VLN) benchmarks typically use discrete or coarse actions and existing UAV Vision-Language-Action (VLA) tasks focus on short, atomic maneuvers. To address this gap in UAV task settings, we introduce \textbf{FLIGHT}, a \textbf{F}ine-grained \textbf{L}ong-horizon \textbf{I}nstruction-\textbf{G}uided benchmark for \textbf{H}ybrid UAV navigation and reasoning \textbf{T}asks, which combines multi-stage instructions with dense 6-DoF trajectory annotations across two dataset splits: Fine-grained VLN and Long-horizon Flow. To endow the UAV agent with the capability of real-time in-flight reasoning over task execution status and mission planning, while simultaneously accommodating high-frequency, real-time precise control, we further propose \textbf{FLIGHT VLA}, an asynchronous architecture that decouples a low-frequency Streaming Pilot Vision-Language Model (VLM) for task-state reasoning from a high-frequency diffusion action model for continuous control, supervised by explicit \textbf{Pilot Reasoning} texts that summarize the current flight state and anticipate the next subgoal. In closed-loop evaluation, FLIGHT VLA consistently surpasses representative VLN and VLA baselines on our FLIGHT benchmarks, achieving stronger multi-stage completion, subgoal adherence, and terminal control. Its trained Streaming Pilot Reasoning VLM further improves UAV video reasoning, validating the effectiveness of our design.
Xiangyi Zheng, Xiangyu Wang, Qinan Liao +6
Colab, Beihang University · National University of Singapore · Meituan
Vision-language navigation (VLN) for UAVs demands grounding free-form instructions into 6-DoF flight under partial observability. While Vision-Language-Action (VLA) models excel at semantic reasoning, they suffer from brittleness due to geometric inconsistency and dynamics mismatch. To address this, we propose ImagineUAV, an imagination-driven framework leveraging cascaded world-action modeling. Instead of direct regression, ImagineUAV employs a latent video diffusion model to generate instruction-conditioned future observations, explicitly imagining environmental evolution, from which 6-DoF motions are inferred via an action extractor. A kinodynamic planner then refines these estimates into collision-free trajectories. Additionally, a step-distilled inference pipeline ensures real-time execution. With only 1.3B parameters, ImagineUAV outperforms prior VLN and VLA baselines on benchmarks and real-world flights, validating the practicality of imagination-driven aerial navigation.
Xuchen Liu, Jiawei Huang, Shihao Xia +3
Pengcheng Laboratory, Shenzhen, Guangdong, China · School of Computer Science and Cyber Engineering, Guangzhou University, Guangzhou, Guangdong, China · Southern University of Science and Technology, Shenzhen, Guangdong, China