Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching
Authors: Guorui Pei, Jinsong Wu, Songyuan Su, Jiaming Qi, Sichao Liu, David Navarro-Alarcon, Bin Liu, Peng Zhou
Organizations: College of Robotics Science and Engineering, Taiyuan University of Technology, Taiyuan, China · Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong SAR, China · Southern University of Science and Technology, Shenzhen, China · School of Advanced Engineering, Great Bay University, Dongguan, Guangdong, China · College of Mechanical and Electrical Engineering, Northeast Forestry University, Harbin, China · Department of Production Engineering, KTH Royal Institute of Technology, Stockholm, Sweden · Kaiyang Laboratory, Chery Automobile Co., Ltd., Wuhu, Anhui, China
Skilled humans can catch fast-moving objects softly by coordinating interception, velocity matching, and follow-through to mitigate impact. Learning such impact-aware catching with reinforcement learning (RL), however, is challenging, as the policy must achieve reliable interception and grasping while regulating the sensitive transition into contact. Moreover, even a capable privileged-state RL teacher may not provide ideal demonstrations for a deployable imitation-learning (IL) student: teacher failures limit task coverage, while small variations in pre-contact motion can produce substantially different impact and grasping outcomes. We characterize this phenomenon through interventional outcome sensitivity and introduce the outcome-sensitive window (OSW) to guide targeted demonstration construction. Building on this formulation, we propose Outcome-Sensitive Motion Search, which learns a task-conditioned manifold of successful OSW motions and performs local geodesic search to refine successful teacher rollouts and repair task conditions where the teacher fails. We then validate candidate motions through complete rollouts under a calibrated IL-student action-error model and retain only successful executions as demonstrations. Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.
Figures & tables
Fig. 1: Motivation for impact-aware dexterous catching. Human and robot examples illustrate interception, velocity matching, and follow-through for impact mitigation.
Fig. 2: Task quantities and success criteria: (a) contact relative speed, peak impact force, and maximum follow-through displacement; (b) conditions for task success.
Fig. 3: Overview of Outcome-Sensitive Motion Search. RL experience flows through event alignment, geodesic search on the OSW motion manifold, student-aware complete-rollout validation, and construction of the final IL training dataset.
Fig. 4: Student-aware candidate selection for (a) successful-rollout refinement and (b) repair.
Fig. 5: Ten object variants and their corresponding impact-mitigating catches. The five families are box, sphere, ellipsoid, cylinder, and capsule, each shown in large red (1) and small yellow (2) variants.
Fig. 6: Representative impact-mitigating catch by the RL teacher. Faded poses show arm–hand and object motion; the red curve traces the object center of mass.
Controller
Catch
Task
vrel
fmax
ℓmax
compl.
succ.
(m/s)
(N)
(m)
Catching-only RL
92.5%
0.4%
5.24
160.1
0.07
RL teacher
90.7%
85.2%
3.60
40.8
0.42
TABLE I: Catching performance with and without impact mitigation.
Demonstration construction
Closed-loop policy performance
Construction / resulting policy
Demo. coverage
Repair rate
Refinement rate
Unperturbed fmax (N)
Task success
Catch compl.
Policy fmax (N)
RL teacher (original)
–
–
–
–
85.2%
90.7%
40.8
Success-only filtering / Naive-IL
85.2%
0%
0%
40.6
83.8% ± 0.6%
89.9% ± 0.5%
41.1 ± 1.0
Unperturbed-only / Normal-IL
94.9%
65.5%
84.2%
35.9
81.7% ± 0.7%
88.0% ± 0.9%
44.2 ± 1.8
Our method / Ours-IL
91.8%
44.3%
59.8%
36.5
90.3% ± 0.8%
93.6% ± 0.4%
37.3 ± 1.2
TABLE II: Demonstration construction and policy performance.
Fig. 8: Impact-mitigating RL teacher catches for a central and three extreme launch directions. Each column shows two views of one rollout; faded arm–hand poses trace catching and follow-through.
Fig. 9: Performance across task conditions. (a) RL teacher task success rate by object family and size; (b) Ours-IL gain over the RL teacher (percentage points); (c) policy task success rate versus absolute lateral target offset.