ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation
Organizations: State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China · Division of Natural and Applied Sciences, Duke Kunshan University, Suzhou, China · Peng Cheng Laboratory, ShenZhen, China
Abstract
Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across different temporal horizons. To address this limitation, we propose ForeFly, a dual-horizon latent world action model that predicts both a proximal future for local continuity and an adaptive route-critical future for long-range guidance. Horizon-specific foresight queries are primed with recent and route-critical visual memories, providing history-aware context for future prediction. To exploit their distinct roles in action generation, we introduce Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance. Experiments on the TravelUAV and UAV-ON benchmarks show that ForeFly consistently outperforms strong baselines across seen and unseen settings, validating the effectiveness of dual-horizon foresight and FGAR learning. The code is available at: https://github.com/kunhuiW/ForeFly
Figures & tables
| Method | Seen | Unseen | ||||
| SR | OSR | SPL | SR | OSR | SPL | |
| AerialVLN Liu et al. (2023) | 4.55 | 16.84 | 4.24 | 3.51 | 20.77 | 3.41 |
| AOA-F Xiao et al. (2025) | 7.87 | 17.98 | 4.21 | 6.45 | 16.09 | 3.98 |
| OpenFly Gao et al. (2026) | 12.29 | 26.20 | 6.83 | 11.82 | 25.71 | 5.64 |
| APEX Zhang et al. (2026a) | 12.36 | 19.10 | 9.13 | 16.13 | 22.58 | 13.03 |
| Navid Zhang et al. (2024) | 14.17 | 30.08 | 7.19 | 10.70 | 28.51 | 5.99 |
| Method | Seen | Unseen | ||||
| SR | OSR | SPL | SR | OSR | SPL | |
| AerialVLN Liu et al. (2023) | 4.55 | 16.84 | 4.24 | 3.51 | 20.77 | 3.41 |
| AOA-F Xiao et al. (2025) | 7.87 | 17.98 | 4.21 | 6.45 | 16.09 | 3.98 |
| OpenFly Gao et al. (2026) | 12.29 | 26.20 | 6.83 | 11.82 | 25.71 | 5.64 |
| APEX Zhang et al. (2026a) | 12.36 | 19.10 | 9.13 | 16.13 | 22.58 | 13.03 |
| Navid Zhang et al. (2024) | 14.17 | 30.08 | 7.19 | 10.70 | 28.51 | 5.99 |
| Method | NE | SR | OSR | SPL |
| MLP Head | 64.41 | 33.08 | 62.91 | 27.99 |
| Direct Dual-Future* | 62.03 | 34.00 | 62.50 | 28.73 |
| FGAR w/o | 54.61 | 38.35 | 68.73 | 32.55 |
| FGAR | 48.74 | 41.75 | 70.25 | 34.75 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.