From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation
Authors: Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, +2 more
Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798 · College of Electronics and Information Engineering, Shanghai Institute of Intelligent Science and Technology, Tongji University, Shanghai 201804, China · School of Aeronautic Science and Engineering, Beihang University, Beijing 100191, China · NTU–VinUni Joint Research Laboratory for Embodied AI and Robotics, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798, and VinUniversity, Hanoi, Vietnam
Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.
Figures & tables
Fig. 2: System overview. (A) Reconfigurable 3DGS-based supervision synthesis. (B) Physical articulated UAM. (C) Synthesized local-skill observations and expert trajectories. (D) Online execution combining Scene Graph grounding and transfer, whole-body VLA prediction, MPAR, and kinodynamic trajectory realization.
Fig. 3: Measured-progress-aligned takeover between consecutive VLA chunks. Nominal-time alignment selects a phase according to elapsed time, whereas MPAR projects the estimated handoff state onto the incoming VLA path. Starting from the active command state at handoff, a C2 bridge reaches a forward feasible join phase, after which MINCO realizes the remaining VLA-path suffix as a continuous trajectory.
Task
Whole-Task VLA
Ours
Water plant
0/10
5/10
Water and discard
0/10
3/10
Discard empty bottle
0/10
6/10
Store rectangular block
0/10
4/10
Store fan-shaped block
0/10
3/10
Overall
0/50
21/50
TABLE I: Cross-site mission performance in simulation. Values denote successful trials over total trials.
Fig. 4: Evaluation of whole-body realization. (A–B) Representative physical local interactions under MPAR. (C) Aggregate success and collision rates for RTC-Direct, the nominal-time ablation, and MPAR. (D–E) Representative simulated storage and watering executions. Green arrows indicate the UAM approach direction, red arrows indicate object motion, and ghosted configurations show the executed whole-body trajectory.
Task/condition
RTC-Direct [ 5 ]
Nominal-time
MPAR
Task-wise success
Grasp Rect
11/40
14/40
12/40
Watering
24/40
27/40
32/40
Storage
19/40
30/40
34/40
Condition-wise success
C0
18/30
21/30
19/30
TABLE II: Task- and condition-wise success under runtime disturbances. Both views summarize the same 120 trials per method (40 per task and 30 per condition); best results are bold. Nominal-time and MPAR denote our ablation and full method, respectively.
Metric
Nominal-time
MPAR
Reduction
Phase P50 [s]
0.510/0.459
0.310/0.273
39.3/40.5%
Phase P95 [s]
0.830/0.831
0.444/0.508
46.6/38.9%
Skipped arc [cm]
8.8/8.7
3.7/3.1
58.0/64.4%
Repeated arc [cm]
12.1/10.3
7.8/5.2
35.5/49.5%
Rejection [%]
17.42
12.81
26.5%
TABLE III: Mechanism-level descriptive comparison of nominal-time alignment and MPAR. Entries report task-balanced C2/C3 aggregates defined in the text; rejection is pooled across C2–C3. Lower is better.
Metric
− G
− R
− H
Full
Stage-level success
Grounding
20/35
35/35
35/35
35/35
Transfer
N/A
27/34
35/35
34/35
Handoff
N/A
26/27
26/35
34/34
Local
N/A
16/26
15/26
26/34
Episode-level outcomes
TABLE IV: Scene Graph physical-interface ablation. − G, − R, and − H remove relational grounding, topology-guided routing, and feasible handoff, respectively. Store, Water, and W → D (Water then Discard) are task-wise mission successes; Mission is their aggregate.
Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are costly, difficult to scale, and hard to repeat under controlled conditions. We present AeroManip-VLA, a scalable benchmark for aerial VLA data generation and policy evaluation. AeroManip-VLA provides a GPU-accelerated simulation framework with low-level payload-aware flight and manipulation control in massively parallel environments. Building on this framework, we combine reusable reinforcement learning policies with expert task rules to automatically generate demonstrations without human teleoperation across diverse objects, environments, and randomized initial conditions. The generated data include basic skills such as grasping and placing, as well as long-horizon tasks that require both navigation and manipulation. We further introduce automated event labeling and trajectory categorization to filter demonstrations. These mechanisms enable fine-grained analysis of task progress, behavioral outcomes, and safety-related failures. Finally, we evaluate a range of imitation learning and VLA baselines across different task settings, revealing their performance characteristics and failure modes. Together, AeroManip-VLA enables scalable aerial manipulation data generation, structured trajectory analysis, and systematic VLA evaluation in simulation prior to real-world deployment.
Rui Huang, Yanlin Mu, Lidong Li +3
National University of Singapore · Beijing Institute of Technology
Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red building directly from camera images. Because VLAs inherit world knowledge from internet-scale pre-training, they have become the dominant framework for learning-based manipulation, with bimanual coordination serving as the most demanding testbed: two arms with 7 degrees of freedom each must move in concert to fold, assemble, and reorient objects. Unmanned aerial robotics faces a structurally similar challenge: a drone must coordinate thrust, attitude, and increasingly gripper commands from visual observations under strict latency and payload constraints. This review covers 183 contributions spanning 2017-2026 and organized along seven dimensions: VLA architectures, training recipes, action representations, bimanual coordination (2022-2026), unmanned aerial vehicle (UAV) navigation and control (2017-2026), language grounding, and cross-cutting concerns including memory and world models. We show that the coordination strategies, training recipes, and action representations developed for bimanual VLAs transfer to unmanned aerial systems and identify fourteen research directions across both domains.
Inkyu Sa, Chanoh Park, Hea-Min Lee +2
Chef Robotics, San Francisco, CA 94103, USA · RovifyLab, Gyeonggi 13840, Republic of Korea · IT Application Research Center, Jeonbuk Regional Branch, Korea Electronics Technology Institute (KETI), Jeonju 54853, Republic of Korea +1
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions. To address the challenges, we present VISTA, a framework that bridges this dual gap through three synergistic components. (i)~UMI-VQA, the first large-scale VQA dataset tailored to wrist-mounted fisheye observations, aligns VLM representations to the distorted visual regime via auxiliary vision-language supervision. (ii)~A systematic physical-validation pipeline performs a data-completeness pre-check and scores each valid trajectory for trajectory continuity, self-collision risk, and execution fidelity before it enters training. (iii)~A two-stage co-training recipe jointly learns vision-language grounding on UMI-VQA and action prediction on validated trajectories. Our experiments empirically show that incorporating UMI-VQA consistently improves downstream policy performance, and that physical-validation scores are strongly predictive of deployment success. On diverse simulation and real-world manipulation tasks, VISTA significantly outperforms strong baselines including π0.5, LingBot-VLA, and Wall-X. We release the physical-validation pipeline, UMI-VQA, validated trajectory data, and the pre-trained model for the community.
Siyuan Yang, Linzheng Guo, Ouyang Lu +10
Institute of AI (TeleAI), China Telecom · University of Science and Technology of China · 4Northwestern Polytechnical University +5