cs.ROOct 6, 2026

WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses

Authors: Thinh D. Le, Son T. Nguyen, Duong Q. Nguyen, Dung D. Le, Ngo Anh Vien, H. Nguyen-Xuan

Organizations: Center for AI Research, VinUniversity, Ho Chi Minh City, Vietnam · College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam · VinRobotics, Hanoi, Vietnam

Abstract

Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

    Jul 7, 2026Inkyu Sa, Chanoh Park, Hea-Min Lee +2Bimanual ManipulationGripper

  2. CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking

    Jul 16, 2026Ruilong Ren, Songsheng Cheng, Yunpeng Zhou +9Aerial Vision-Language NavigationUnmanned Aerial Vehicle Detection

  3. Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations

    Jul 3, 2026Van Huyen Dang, Kabilesh Rajendran, Erdi Sayar +1Autonomous DronesUnmanned Aerial Vehicles