Whole-Body Aerial Grasping and Lifting via Partial Visual Observations
Authors: Jiaye Jin, Rui Jin, Xinhang Xu, Haotian Jin, Ruiyang Liu, Yi Wang, Jiayan Zhao, Kun Cao, +1 more
Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798 · College of Electronics and Information Engineering, Shanghai Institute of Intelligent Science and Technology, Tongji University, Shanghai 201804, China · NTU–VinUni Joint Research Laboratory for Embodied AI and Robotics, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798, and VinUniversity, Hanoi, Vietnam
Aerial grasp-and-lift tasks require whole-body coordination across approach, acquisition, and lifting under partial target observations. Early approach failures can limit exposure to later task stages during training, while changing visibility complicates alignment and closure timing during execution. We present a recurrent teacher-student framework that learns a single policy in simulation to jointly command flight, arm motion, and gripper closure without an explicit task-phase input. A privileged teacher learns through reinforcement learning with a critical-state curriculum that exposes acquisition and lifting states before connecting them to normal approach trajectories. Its behavior is distilled into a recurrent visual student that replaces privileged target states with dual-view point clouds and proprioception, integrating observation history for closed-loop control. A dedicated closure objective supervises closure timing from sustained model-defined readiness sequences. Training and primary evaluation use a simulated acquisition-and-payload model with condition-triggered latching, virtual attachment, and wrench-based payload loading for short-distance lifting. Across 8,996 completed simulation episodes under this model, the frozen student achieves full-task success rates of 99.97%, 97.14%, and 95.84% under nominal, physics/control-randomized, and additional camera-randomized conditions, respectively. The nominal latch-count-weighted mean of per-seed 90th-percentile alignment errors at acquisition is 8.12 mm.
Figures & tables
Fig. 1: Aerial acquisition-and-lift scenario. Left: simulated approach, modeled acquisition, and lifting. Right: photographs of the physical platform and task scenario, shown for illustration only.
Fig. 2: Aerial manipulation platform. Left: physical prototype and camera placement. Right: simulation model used for training and evaluation.
Fig. 3: Overview of the recurrent teacher–student framework. The privileged teacher is distilled into a recurrent visual student conditioned on dual-view point clouds and proprioception.
Parameter
Setting
Thrust-to-weight ratio
[2.0,3.2]
Thrust delay coefficient α
[0.22,0.50]
Body-rate delay coefficient α
[0.33,0.63]
Action delay
1–3 control steps
Thrust noise std.
5%
IMU noise
Linear/angular velocity, gravity
TABLE I: Domain randomization settings.
Training schedule
Nominal (%)
DR (%)
Direct full task
33.33
32.59
Simple two stage
0.00
0.00
Full curriculum
100.00
99.90
TABLE II: Full-task success under the acquisition-and-payload model for alternative training schedules. Direct and simple two-stage results are averaged across three independently trained policies; each policy is evaluated with three rollout seeds. The full-curriculum result is from the selected reference policy, evaluated with the same three rollout seeds.
Fig. 4: Full-task success versus cumulative environment timesteps for three training schedules under (a) nominal and (b) physics/control domain-randomized conditions. Curves connect unsmoothed, equal-weight means across three training seeds; checkpoints missing any seed are omitted, breaking the lines.
Policy
Nominal (%)
DR (%)
+ Camera (%)
Privileged teacher
100.00
99.90
—
BC-only student
0.00
0.00
—
Final student
99.97
97.14
95.84
TABLE III: Full-task success under the configured acquisition-and-payload model. Each cell pools three evaluation seeds for one checkpoint; dashes denote unevaluated conditions.
Fig. 5: Simulated acquisition and lifting sequence. (a) End-effector alignment error. (b) Finger positions and gripper command (right axis). (c) Height gain; dashed line: 150-mm lift threshold. (d) Relative end-effector speed. Shading indicates execution phases: approach ( t0 ), acquisition ( t1 ), and lift ( t2 ), and is used for visualization only.
Fig. 6: Visual-input and payload sensitivity under the acquisition-and-payload model. (a) Camera removal from the frozen student without retraining; results pool three rollout seeds. Wrist-only uses D405; body-only uses D450. Combined DR includes physics/control and camera randomization. (b) Open circles: single-seed mass sweeps; filled squares: three-seed confirmations. Error bars: 95% Wilson episode-level intervals. Dashed and dotted lines mark pooled and per-seed acceptance thresholds, respectively.
Fig. 7: MuJoCo evaluation setup. Left: candidate regions for randomized initialization of aircraft (purple) and object (orange) positions; sampled configurations are screened for initial target visibility. Right: the three test objects.
Fig. 8: Qualitative simulated acquisition-and-lift sequences for (a) a sugar box, (b) a Coke can, and (c) a screwdriver. Translucent overlays depict intermediate robot poses.
Object
Acquisition–lift
Full task
Failures
Sugar box
29/30
25/30
1 contact; 4 hold
Coke can
28/30
28/30
2 contact
Screwdriver
30/30
30/30
None
TABLE IV: Acquisition-and-lift and full-task success in MuJoCo (30 trials per object).
Aerial grasping is a remarkable capability exhibited by predatory birds, allowing them to capture prey through highly coordinated maneuvers in flight. Inspired by this capability, researchers have developed various formulations to reproduce such maneuvers through trajectory optimization. However, two limitations remain in practice. First, the resulting optimization problem is highly nonconvex and sensitive to initialization, making high-quality solutions difficult to obtain under a limited computational budget. Second, prescribed numerical objectives are human-designed abstractions that describe successful grasping through a limited set of mathematically tractable quantities and may not fully capture what determines task success. We investigate how learning can address these limitations within an analytical planner. Accordingly, a trajectory prior is first learned from optimized motions and then evolved through a CEM-based process that evaluates sampled initializations with the deployed optimizer and retains favorable ones as new supervision. An Execution-Aware Critic learns from contact, lift, and completion outcomes to assess whether the optimized trajectories are likely to succeed in physical execution. Its frozen energy can further serve as a differentiable grasping cost, allowing execution data to directly shape trajectory generation. Simulation and real-world experiments demonstrate improved optimization reliability, trajectory consistency, and grasping performance.
Weiliang Deng, Zhengyang Dang, Yao Mu +1
School of Intelligent Systems Engineering, Sun Yat-sen University, Guangzhou 510275, China · AI Institute, School of Computer Science, Shanghai Jiao Tong University, Shanghai 200240, China · Differential Robotics Technology Co., Ltd., Hangzhou, China
Reliable aerial grasping in cluttered environments remains challenging due to occlusions and collision risks. Existing aerial manipulation pipelines largely rely on centroid-based grasping and lack integration between the grasp pose generation models, active exploration, and language-level task specification, resulting in the absence of a complete end-to-end system. In this work, we present an integrated pipeline for reliable aerial grasping in cluttered environments. Given a scene and a language instruction, the system identifies the target object and actively explores it to gain better views of the object. During exploration, a grasp generation network predicts multiple 6-DoF grasp candidates for each view. Each candidate is evaluated using a collision-aware feasibility framework, and the overall best grasp is selected and executed using standard trajectory generation and control methods. Experiments in cluttered real-world scenarios demonstrate robust and reliable grasp execution, highlighting the effectiveness of combining active perception with feasibility-aware grasp selection for aerial manipulation.
Fast dexterous grasping while a mobile base remains in motion requires coordinated whole-body control and rapid adaptation to physical contact. We propose FastGrasp, a two-stage learning framework that integrates grasp guidance, whole-body control, and tactile feedback. First, a pretrained conditional variational autoencoder generates diverse grasp candidates from object point clouds. Candidates are filtered by approach direction and supporting-surface constraints, then ranked using an envelopment-based criterion combining grasp width and depth coverage measures. Second, a reinforcement learning policy jointly controls the mobile base, arm, and dexterous hand, using the selected grasp as guidance and tactile observations for online grasp adjustment. The policy is trained with domain randomization and deployed with command filtering and tactile-triggered grasp tightening. Simulation experiments show higher grasp success rates than the evaluated baselines under full and partial point-cloud observations, while real-world experiments demonstrate sim-to-real transfer across diverse object geometries.