Kinematic Nonlinear Spatio-Temporal Trajectory Warping for Contact-Rich Dexterous Manipulation Demonstrations
Authors: Hyojae Park, Arjun S. Lakshmipathy, Nancy S. Pollard
Organizations: Computer Science Department and the 1Robotics Institute at Carnegie Mellon University, Pittsburgh, USA. · Robotics Institute at Carnegie Mellon University, Pittsburgh, USA.
We present a straightforward but effective method for repurposing existing contact-rich dexterous manipulation demonstrations. Starting from inputs of hand and object trajectories, our method outputs high-quality nonlinear trajectory warps that account for intermediate waypoints, environmental barriers, temporal shifts, and varied start/end configurations. Foundational to our method is the utilization of contact distributions, which we show allows us to reliably compute complex and high-dimensional dexterous hand trajectories following a simple object-centric warp specification pipeline. We evaluate our method across 12 variations sourced from 4 demonstrations in a publicly available dataset of human hand motion data, perform baseline comparisons, and demonstrate generalization of our approach to different manipulators. Results and code will be made available on publication.
Figures & tables
Fig. 1 : Our method warps hand and object trajectories from existing contact-rich demonstrations to fit spatial and temporal waypoints, environmental barriers, and new start/end positions. Approximate mappings between the original and warped trajectories are color-coded for visualization.
Fig. 2 : System overview. Given a motion demonstration, our method first spatially warps the object trajectory to intersect user-specified waypoints and avoid scene barriers. It then applies a time-warp to ensure the trajectory satisfies the time constraints of each waypoint. Finally, an inverse kinematics (IK) optimization recovers the hand motion by tracking the warped object trajectory, subject to hand–barrier collision constraints, using per-frame corresponding contact distributions to yield the complete retargeted demonstration.
Figure 3
Trajectory
# of waypoints
# of barriers
pstart′ provided
pend′ provided
cube_messy
0
7
✓
✓
cube_messy2
4
8
✗
✓
fryingpan_change_tops
1
0
✓
✓
fryingpan_island
3
0
✓
✓
fryingpan_onto_shelf
4
3
✓
✓
mug_pass_wall
2
1
✗
✗
TABLE I : Evaluation trials and input specifications
Trajectory
Init Mean Dist
10 mm
5 mm
3.5 mm
cube_messy
150.7
148.5
150.7
150.8
cube_messy2
150.7
149.0
150.0
149.7
fryingpan_change_tops
164.5
165.7
164.2
165.7
fryingpan_island
164.5
156.6
159.8
162.6
fryingpan_onto_shelf
164.5
165.0
166.1
166.4
mug_pass_wall
145.4
143.3
145.3
145.5
TABLE II : DWO across different optimization thresholds. Values are in millimeters (mm).
Trajectory
Init Cont.
10 mm
5 mm
3.5 mm
Drift
cube_messy
4.6
9.2
4.9
4.3
-0.24
cube_messy2
4.6
9.4
6.0
5.5
0.12
fryingpan_change_tops
5.5
7.9
8.5
5.7
-0.37
fryingpan_island
5.5
9.2
6.9
6.1
-0.70
fryingpan_onto_shelf
5.5
10.0
6.9
6.1
-0.29
mug_pass_wall
8.5
10.6
6.6
6.1
-0.79
TABLE III : DC across varied optimization thresholds in mm.
Trajectory
# of Timesteps
10 mm
5 mm
3.5 mm
cube_messy
673
74.63
196.93
464.27
cube_messy2
673
96.64
226.91
481.55
fryingpan_change_tops
648
28.39
80.81
206.03
fryingpan_island
648
37.05
118.83
214.40
fryingpan_onto_shelf
648
35.59
110.07
212.03
mug_pass_wall
544
61.79
115.33
174.64
TABLE IV : Runtime across different optimization thresholds. Non-timestep values are in seconds.
PLR Hand
Trajectory
10 mm
5 mm
3.5 mm
PLR Object
Timewarp
cube_messy
0.999
1.002
0.989
1.005
0.075
cube_messy2
1.358
1.378
1.389
2.024
0.085
fryingpan_change_tops
1.019
0.945
0.933
0.897
0.276
fryingpan_island
1.519
1.522
1.527
2.165
0.072
fryingpan_onto_shelf
1.361
1.314
1.321
1.733
0.075
TABLE V : PLRs and timewarp discrepencies across different optimization thresholds.
Fig. 4 : Contact correspondence robustness test on the teapot_pour_cups_wall scene. Both curves maintain correspondence quality up until 50% ( Dc stays near 5.0 mm).
Fig. 5 : An example warping using the Franka arm. The hand must move the box through a series of waypoints.
Fig. 6 : An example Allegro hand warping where (a) the input apple handoff trajectory is (b) altered to pass through waypoints (red) and terminate on an elevated plate. (c) The timewarp visualization illustrates where frames of the original trajectory fall in the warped sequence, with black portions indicating segments which must be interpolated.
Fig. 7 : Comparison of the baseline and our hand warping methods. The yellow region marks the in-contact segment. Light purple lines show the timewarp mapping from input to output timesteps. Four timesteps are visualized for both warping methods (cyan, orange, gray, brown). The baseline method degrades at the orange and gray timesteps, where ground-truth transformation is not available. Our contact-based method continues to maintain a proper grasp across all timesteps.
Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.
Reinforcement learning explores effectively in domains such as Atari games, navigation, and locomotion, where novelty over states or dynamics is a sufficient signal. In contrast, dexterous manipulation requires rich physical hand--object interactions, but existing methods often suffer from unstable contact-based novelty signals, inefficient distance novelty signals, or reliance on task-specific priors. We propose ContactExplorer, a general exploration method for dexterous manipulation tasks. ContactExplorer represents contact as the intersection between object surface points and hand keypoints, encouraging dexterous hands to discover diverse and novel contact patterns, namely which fingers contact which object regions. It maintains a contact counter conditioned on discretized object states obtained via learned hash codes. This counter is leveraged in two complementary ways: (1) a count-based contact coverage reward that promotes exploration of novel contact patterns, and (2) an energy-based reaching reward that guides the agent toward under-explored contact regions. We evaluate ContactExplorer on seven contact-rich manipulation tasks and five dexterous hand embodiments. Experimental results show that ContactExplorer substantially improves sample efficiency and success rates over existing exploration methods, that it reduces the need for task-specific priors, and that it remains effective across hand embodiments and transfers to the real world. Project page is https://contact-explorer.github.io.
Zixuan Liu, Ruoyi Qiao, Chenrui Tie +5
School of Computing, National University of Singapore · RoboScience
Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.
Zihao Yang, Chengyuan Liu, Yu Zhou +6
DexGEM Lab · Shanghai Jiao Tong University · Tongji University +1