Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching
Authors: Guorui Pei, Jinsong Wu, Songyuan Su, Jiaming Qi, Sichao Liu, David Navarro-Alarcon, Bin Liu, Peng Zhou
Organizations: College of Robotics Science and Engineering, Taiyuan University of Technology, Taiyuan, China · Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong SAR, China · Southern University of Science and Technology, Shenzhen, China · School of Advanced Engineering, Great Bay University, Dongguan, Guangdong, China · College of Mechanical and Electrical Engineering, Northeast Forestry University, Harbin, China · Department of Production Engineering, KTH Royal Institute of Technology, Stockholm, Sweden · Kaiyang Laboratory, Chery Automobile Co., Ltd., Wuhu, Anhui, China
Skilled humans can catch fast-moving objects softly by coordinating interception, velocity matching, and follow-through to mitigate impact. Learning such impact-aware catching with reinforcement learning (RL), however, is challenging, as the policy must achieve reliable interception and grasping while regulating the sensitive transition into contact. Moreover, even a capable privileged-state RL teacher may not provide ideal demonstrations for a deployable imitation-learning (IL) student: teacher failures limit task coverage, while small variations in pre-contact motion can produce substantially different impact and grasping outcomes. We characterize this phenomenon through interventional outcome sensitivity and introduce the outcome-sensitive window (OSW) to guide targeted demonstration construction. Building on this formulation, we propose Outcome-Sensitive Motion Search, which learns a task-conditioned manifold of successful OSW motions and performs local geodesic search to refine successful teacher rollouts and repair task conditions where the teacher fails. We then validate candidate motions through complete rollouts under a calibrated IL-student action-error model and retain only successful executions as demonstrations. Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.
Figures & tables
Fig. 1: Motivation for impact-aware dexterous catching. Human and robot examples illustrate interception, velocity matching, and follow-through for impact mitigation.
Fig. 2: Task quantities and success criteria: (a) contact relative speed, peak impact force, and maximum follow-through displacement; (b) conditions for task success.
Fig. 3: Overview of Outcome-Sensitive Motion Search. RL experience flows through event alignment, geodesic search on the OSW motion manifold, student-aware complete-rollout validation, and construction of the final IL training dataset.
Fig. 4: Student-aware candidate selection for (a) successful-rollout refinement and (b) repair.
Fig. 5: Ten object variants and their corresponding impact-mitigating catches. The five families are box, sphere, ellipsoid, cylinder, and capsule, each shown in large red (1) and small yellow (2) variants.
Fig. 6: Representative impact-mitigating catch by the RL teacher. Faded poses show arm–hand and object motion; the red curve traces the object center of mass.
Controller
Catch
Task
vrel
fmax
ℓmax
compl.
succ.
(m/s)
(N)
(m)
Catching-only RL
92.5%
0.4%
5.24
160.1
0.07
RL teacher
90.7%
85.2%
3.60
40.8
0.42
TABLE I: Catching performance with and without impact mitigation.
Demonstration construction
Closed-loop policy performance
Construction / resulting policy
Demo. coverage
Repair rate
Refinement rate
Unperturbed fmax (N)
Task success
Catch compl.
Policy fmax (N)
RL teacher (original)
–
–
–
–
85.2%
90.7%
40.8
Success-only filtering / Naive-IL
85.2%
0%
0%
40.6
83.8% ± 0.6%
89.9% ± 0.5%
41.1 ± 1.0
Unperturbed-only / Normal-IL
94.9%
65.5%
84.2%
35.9
81.7% ± 0.7%
88.0% ± 0.9%
44.2 ± 1.8
Our method / Ours-IL
91.8%
44.3%
59.8%
36.5
90.3% ± 0.8%
93.6% ± 0.4%
37.3 ± 1.2
TABLE II: Demonstration construction and policy performance.
Fig. 8: Impact-mitigating RL teacher catches for a central and three extreme launch directions. Each column shows two views of one rollout; faded arm–hand poses trace catching and follow-through.
Fig. 9: Performance across task conditions. (a) RL teacher task success rate by object family and size; (b) Ours-IL gain over the RL teacher (percentage points); (c) policy task success rate versus absolute lateral target offset.
Fast catching of free-flying objects is difficult because of short reaction time, impact uncertainty, and kinodynamic constraints. We use reinforcement learning in simulation to collect successful catching trajectories and learn a low-dimensional kinodynamic trajectory manifold. At run time, the estimated object initial state is mapped directly to a reference catching trajectory without online nonlinear optimization. The trajectory is tracked with compliant control near contact for improved impact absorption and capture stability.
Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, objects, or instructions, but their performance across task execution speeds is less often examined. This leaves open how much temporal robustness a learner retains relative to the expert it imitates. We compare an expert and learner under the same task conditions, initial-condition draws, and speedup factors. We instantiate the evaluation in ParcelStow, a contact-rich task in which the robot acquires, reorients, and inserts a parcel. The demonstrations span the speedup range for the manipulation phases after parcel acquisition. A scripted expert and an Action Chunking with Transformers (ACT) policy trained from the expert's demonstrations both achieve 100 percent task success at nominal speed. Their success rates diverge within the demonstrated range: at its maximum, expert success is 84 percent and ACT success is 53 percent. Two ACT policies with different parameter initializations show similar degradation, decreasing by 34 and 48 percentage points from nominal speed to the maximum demonstrated speed, compared with 16 points for the expert. Stage-level analysis shows that 35 of ACT's 47 failures at the maximum demonstrated speed are insertion misalignments. Under the relative-motion handoff, every ACT acquisition retains the parcel through reorientation and transfer in free space, but only 64 percent complete the overall task, compared with 95 percent after expert acquisition. Across all evaluated policies and speeds, none of the 414 acquisitions without force closure completes the task. Equal nominal task success therefore does not imply preservation of expert performance across execution speeds. Code, data, and evaluation scripts are available at https://github.com/coenwerem/parcelstow.
Clinton Enwerem, John S. Baras, Calin Belta
Institute for Systems Research University of Maryland College Park, MD, U.S.A.
Real-world learning for dexterous hands remains brittle because high-dimensional hand actions amplify imitation errors and make reinforcement-learning exploration prone to contact-breaking motion. While combining imitation learning (IL) with online reinforcement learning (RL) can reduce manual supervision, unconstrained exploration in raw hand-action spaces is sample-inefficient and risky for physical hardware. We introduce a latent motion prior module (\prior{}) that maps recent hand-action histories to a compact, history-conditioned latent prior and decodes continuous latent commands into executable high-dimensional hand targets. Built on this prior, \method{} is a three-stage real-world dexterous learning framework: it pretrains \prior{} from demonstrations, trains a visuomotor policy that predicts native arm commands and latent hand-action offsets, and improves the policy with online residual RL in the same latent hand-action space. This shared, decodable interface lets residual exploration make local corrections near demonstrated, contact-consistent hand motions rather than perturbing every finger joint independently. We evaluate \method{} on four real-robot dexterous manipulation tasks against raw, linear, and discrete hand-action interfaces. Starting from small task-specific demonstration sets, \method{} achieves a 56.25% average IL success rate and raises it to 98.75% after online RL, reaching 100% final success on three tasks and 95% on the remaining task.
Xinye Yang, Zhiyuan Ma, Hongze Yu +5
Fudan University · PsiBot · Zhongguancun Academy +2