WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses
Organizations: Center for AI Research, VinUniversity, Ho Chi Minh City, Vietnam · College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam · VinRobotics, Hanoi, Vietnam
Abstract
Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.
Figures & tables
| Property | Value |
| Total episodes | 507 |
| Non-terminal transitions | 8,504 |
| train / val | 7,276 / 1,228 |
| Train / val episodes | 431 / 76 |
| Frame rate | 1 FPS |
| Image resolution | px |
| Component | Mean | Std | Min | Max |
| (m) | ||||
| (m) | ||||
| (m) | ||||
| (rad) |
| Dataset | Plat- form | Lang. instr. | Cont. acts. | Src. | Episodes |
| UAV123 [ 33 ] | UAV | ✓ | Real | 150 | |
| AerialVLN [ 30 ] | UAV | ✓ | Sim | 8,446 | |
| AirNav [ 12 ] | UAV | ✓ | Real | 35,000 | |
| UAV-VLA [ 41 ] | UAV | ✓ | Real | 100k | |
| RaceVLA [ 42 ] | UAV | ✓ | ✓ | Real | 200 |
| CognitiveDrone [ 34 ] | UAV | ✓ | ✓ | Sim | 8,000 |
| Model | Parameters | Action Representation | Vision–Language Backbone | Adaptation |
| OpenVLA-7B | 7.6B total 49.3M trainable | Autoregressive policy (256 bins per dimension) | Prismatic SigLIP + DINOv2 Llama-2 7B | LoRA |
| GR00T N1.7 | 3.1B total 1.62B trainable | Diffusion policy Horizon = 4 | Cosmos-Reason2-2B | Fine-tuning |
| /OpenPI | 3.3B total 145M trainable | Flow matching Horizon = 4 | PaliGemma-3B | LoRA |
| SmolVLA | 450M total 450M trainable | Flow matching Horizon = 4 | SmolVLM-2 + Action Expert | Full fine-tuning |
| Model | (m) | (m) | (m) | (rad) |
| MAE | ||||
| Naive (constant-mean) | 0.590 | 0.242 | 0.003 | 0.248 |
| OpenVLA-7B [ 27 ] | 0.620 | 0.368 | 0.005 | 0.409 |
| GR00T N1.7 [ 38 ] † | 5.728 | 3.867 | 0.637 | 0.555 |
| /OpenPI [ 5 ] | 0.413 | 0.244 | 0.003 | 0.220 |
| SmolVLA [ 43 ] | 0.468 | 0.264 | 0.003 | 0.243 |
| Model | (m) | (m) | (m) | (rad) |
| MAE ( , except ) | ||||
| Naive | ||||
| 4.80 0.09 | 2.70 0.06 | 0.037 | 2.62 0.06 | |
| SmolVLA | 5.33 0.09 | 2.95 0.06 | 0.038 | 2.68 0.06 |
| Pearson | ||||
| Rate | Model | Params | Lat. (ms) | p95 | Hz |
| 1 FPS | SmolVLA | B | |||
| /OpenPI | B | ||||
| GR00T N1.7 | B | ||||
| OpenVLA-7B | B | ||||
| 10 FPS | SmolVLA | B | |||
| /OpenPI | B |
| Rate | Model | ADE | FDE |
| 1 FPS | /OpenPI | ||
| SmolVLA | |||
| OpenVLA-7B | |||
| GR00T N1.7 | |||
| 10 FPS | /OpenPI | ||
| SmolVLA |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| OpenVLA | GR00T | SmolVLA | ||
| Adaptation | LoRA | proj.+DiT | LoRA | full |
| LoRA | expert head | |||
| Optimiser | AdamW | AdamW | AdamW | AdamW |
| Learning rate | ||||
| Steps / epochs | ep | st | st | st |
| Batch size |
| Model | ||||
| openvla (train) | 0.472 | 0.297 | 0.004 | 0.358 |
| openvla (val) | 0.620 | 0.368 | 0.005 | 0.409 |
| gap | +0.149 | +0.070 | +0.001 | +0.052 |
| smolvla (train) | 0.083 | 0.050 | 0.002 | 0.063 |
| smolvla (val) | 0.468 | 0.264 | 0.003 | 0.243 |
| gap | +0.385 | +0.214 | +0.002 | +0.180 |
| Verb | OpenVLA | GR00T | SmolVLA | |
| accompany | 0.300 (56) | 2.594 (53) | 0.221 (56) | 0.197 (56) |
| approach | 0.258 (229) | 2.595 (289) | 0.180 (229) | 0.199 (229) |
| fly after | 0.665 (19) | 3.317 (14) | 0.255 (19) | 0.245 (19) |
| follow | 0.310 (108) | 3.349 (133) | 0.202 (108) | 0.233 (108) |
| go after | 0.554 (46) | 2.719 (62) | 0.264 (46) | 0.297 (46) |
| move closer | 0.382 (52) | 2.383 (45) | 0.187 (52) | 0.190 (52) |