From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation
Authors: Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, +2 more
Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798 · College of Electronics and Information Engineering, Shanghai Institute of Intelligent Science and Technology, Tongji University, Shanghai 201804, China · School of Aeronautic Science and Engineering, Beihang University, Beijing 100191, China · NTU–VinUni Joint Research Laboratory for Embodied AI and Robotics, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798, and VinUniversity, Hanoi, Vietnam
Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.
Figures & tables
Fig. 2: System overview. (A) Reconfigurable 3DGS-based supervision synthesis. (B) Physical articulated UAM. (C) Synthesized local-skill observations and expert trajectories. (D) Online execution combining Scene Graph grounding and transfer, whole-body VLA prediction, MPAR, and kinodynamic trajectory realization.
Fig. 3: Measured-progress-aligned takeover between consecutive VLA chunks. Nominal-time alignment selects a phase according to elapsed time, whereas MPAR projects the estimated handoff state onto the incoming VLA path. Starting from the active command state at handoff, a C2 bridge reaches a forward feasible join phase, after which MINCO realizes the remaining VLA-path suffix as a continuous trajectory.
Task
Whole-Task VLA
Ours
Water plant
0/10
5/10
Water and discard
0/10
3/10
Discard empty bottle
0/10
6/10
Store rectangular block
0/10
4/10
Store fan-shaped block
0/10
3/10
Overall
0/50
21/50
TABLE I: Cross-site mission performance in simulation. Values denote successful trials over total trials.
Fig. 4: Evaluation of whole-body realization. (A–B) Representative physical local interactions under MPAR. (C) Aggregate success and collision rates for RTC-Direct, the nominal-time ablation, and MPAR. (D–E) Representative simulated storage and watering executions. Green arrows indicate the UAM approach direction, red arrows indicate object motion, and ghosted configurations show the executed whole-body trajectory.
Task/condition
RTC-Direct [ 5 ]
Nominal-time
MPAR
Task-wise success
Grasp Rect
11/40
14/40
12/40
Watering
24/40
27/40
32/40
Storage
19/40
30/40
34/40
Condition-wise success
C0
18/30
21/30
19/30
TABLE II: Task- and condition-wise success under runtime disturbances. Both views summarize the same 120 trials per method (40 per task and 30 per condition); best results are bold. Nominal-time and MPAR denote our ablation and full method, respectively.
Metric
Nominal-time
MPAR
Reduction
Phase P50 [s]
0.510/0.459
0.310/0.273
39.3/40.5%
Phase P95 [s]
0.830/0.831
0.444/0.508
46.6/38.9%
Skipped arc [cm]
8.8/8.7
3.7/3.1
58.0/64.4%
Repeated arc [cm]
12.1/10.3
7.8/5.2
35.5/49.5%
Rejection [%]
17.42
12.81
26.5%
TABLE III: Mechanism-level descriptive comparison of nominal-time alignment and MPAR. Entries report task-balanced C2/C3 aggregates defined in the text; rejection is pooled across C2–C3. Lower is better.
Metric
− G
− R
− H
Full
Stage-level success
Grounding
20/35
35/35
35/35
35/35
Transfer
N/A
27/34
35/35
34/35
Handoff
N/A
26/27
26/35
34/34
Local
N/A
16/26
15/26
26/34
Episode-level outcomes
TABLE IV: Scene Graph physical-interface ablation. − G, − R, and − H remove relational grounding, topology-guided routing, and feasible handoff, respectively. Store, Water, and W → D (Water then Discard) are task-wise mission successes; Mission is their aggregate.
Chef Robotics, San Francisco, CA 94103, USA · RovifyLab, Gyeonggi 13840, Republic of Korea · IT Application Research Center, Jeonbuk Regional Branch, Korea Electronics Technology Institute (KETI), Jeonju 54853, Republic of Korea +1