Authors: Meizhong Wang, Kun Cao, Ruiqi Ni, Lihua Xie, Yiguang Hong
Organizations: Tongji University · Shanghai Research Institute for Intelligent Autonomous Systems · Purdue University · Nanyang Technological University, Singapore
Human demonstrations offer rich examples of precise dexterous manipulation and a promising source of robot training data. However, high-fidelity reproduction of demonstrated motions and hand-object interactions across robot embodiments remains challenging under physical constraints. We present DexForge, a differentiable physics-grounded framework for converting human video demonstrations into high-fidelity robot trajectories. We reconstruct spherical-Gaussian object models and hand-object motion from visual observations, then build a differentiable simulator combining efficient Gaussian collision detection with existing differentiable dynamics. Based on this simulator, DexForge combines contact-aware kinematic retargeting with force-aware dynamics retargeting: robot-adapted stable contacts guide kinematic reference construction and subsequent gradient-based control refinement for precise physical motion reproduction. Experiments on 130 DexYCB and HOT3D demonstrations across seven dexterous hands show success-rate gains of approximately 35-53 percentage points over the baseline, with object position and orientation tracking errors on successful trajectories reduced by approximately 34-67% and 71-78%, respectively. Further experiments demonstrate open-loop transfer to MuJoCo and real-robot execution. Our project page is available at https://wmz1226.github.io/DexForge/
Figures & tables
Figure 1: Overview of DexForge. (a) Hand and object reconstruction builds spherical Gaussian object models under RGB-D supervision and recovers hand–object motion from demonstration videos. (b) Contact-aware kinematic retargeting refines demonstrated contacts for robot reachability and contact stability, generating penetration-free references that track them. (c) Force-aware dynamics retargeting optimizes controls through simulation gradients in ComFree-GS, which combines efficient differentiable Gaussian collision detection with ComFree dynamics. System identification further reduces target-system mismatch to support high-fidelity physical reproduction.
Figure 2: Quantitative results on all 20 objects in DexYCB (top); front-view rendering and distance-field visualization for a representative object (bottom).
Figure 3: Dexterous hands and diverse objects used in evaluation. DexForge is evaluated on diverse objects and seven robotic hands that differ in size, degrees of freedom, and finger count.
Figure 4: Visualization of selected kinematic retargeting results from Contact-aware and IK.
Figure 5: Qualitative results of force-aware dynamics retargeting. (a) Motion reproduction by our method and baselines relative to the MANO reference. (b) Retargeting across seven robotic hands.
Figure 6: Sim-to-real open-loop transfer. Controls optimized in ComFree-GS are directly replayed on a Unitree G1 equipped with a LEAP Hand v1.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
Object pose
Hand accuracy
Geometric consistency
Translation (mm)
Rotation ( ∘ )
World-frame MPJPE (mm)
Wrist-aligned MPJPE (mm)
Penetration depth (mm)
Penetrating frames (%)
DexYCB
GT
0.00±0.00
0.00±0.00
0.00±0.00
0.00±0.00
3.15±1.79
59.5±13.6
Ours
0.40±0.16
0.46±0.38
6.42±1.18
4.82±0.86
0.24±0.15
3.9±8.6
HOT3D
GT
0.00±0.00
0.00±0.00
0.00±0.00
0.00±0.00
6.13±3.04
97.8±7.3
Ours
12.23±3.11
4.83±0.78
23.47±2.34
9.21±1.21
1.38±0.21
7.8±4.7
Appendix
Table 5: Hand–object reconstruction results on real multi-view videos from DexYCB and HOT3D (mean ± SD across sequences). GT denotes dataset annotations. Lower is better for all metrics.
Figure 7: Visualization of the reconstructed hand motion and the estimated object pose on a DexYCB sequence, overlaid on the original video from one camera view. Columns show four evenly spaced frames in temporal order. Top: the reconstructed 21 hand keypoints projected into the image. Bottom: the reconstructed MANO hand mesh and the estimated object pose, visualized as the green 3D bounding box of the object’s Gaussian model.
Figure 8: Forward collision-query time with and without BVH for 1–64 parallel environments on an NVIDIA RTX 4090D GPU.
Figure 9: Performance–runtime comparison between DexForge and Contact-aware+SPIDER. Object position (a) and rotation (b) tracking errors are plotted against optimization time, with runtime varied through the optimization iteration limit of each method. Curves show the mean over three runs and shaded regions indicate the standard deviation. Under the same runtime budget, DexForge consistently achieves lower tracking errors.
Figure 10: Sensitivity to control-knot spacing under a matched 200s optimization budget and H=4 control intervals. Errors are computed on a fixed subset of ablation sequences successfully retargeted by both methods under their default configurations; at other spacings, all runs on this subset are included regardless of success. Panels show object position (a) and rotation (b) tracking errors. Curves show mean errors over 3 runs, with shaded regions indicating ±1 standard deviation. Contact-aware+SPIDER is largely insensitive to knot spacing, whereas DexForge achieves lower errors at finer spacings but degrades as the spacing and prediction-window duration increase, consistent with greater difficulty in gradient propagation through longer rollouts.
Human hand-object demonstrations provide a scalable source of data for dexterous robot learning, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based methods typically optimize each demonstration independently, leading to either limited success under finite simulation budgets or training costs that grow with dataset size. We introduce FlashDexRetarget, an RL framework for multi-reference dexterous retargeting. We formulate retargeting as multi-reference tracking, jointly learning a single policy across many demonstrations with off-policy RL and geometric supervision of the demonstrated interactions. This shared training formulation amortizes optimization across references while enabling the policy to track diverse hand-object interactions. On a 50-motion benchmark from TACO, OakInk2, and HOT3D using XHand and Sharpa Wave Hand as target embodiments, FlashDexRetarget retargets 90% of demonstrations using about 30 GPU-hours, compared with about 46% at about 3,000 GPU-hours for CHORD. This corresponds to about 100 times lower training compute and a 44-percentage-point improvement in retargeting success. Ablations examine the key design choices, while experiments with up to 1,000 motions and real-world replay further demonstrate the scalability and practical applicability of our method.
Learning dexterous manipulation from human-object interaction (HOI) data offers a scalable alternative to robot teleoperation, but HOI demonstrations are typically sparse and purely kinematic, making direct retargeting unreliable under embodiment mismatch and contact-rich dynamics. We present DexSynRefine, a coupled framework that treats HOI data as structured motion priors rather than executable robot actions. DexSynRefine first synthesizes hand-object trajectories conditioned on the task and initial object state using HOI Motion Manifold Flow Primitives (HOI-MMFP), a motion prior for coupled hand-object motion. It then physically grounds them with task-space residual reinforcement learning and adapts execution by inferring missing contact-dynamics context from proprioceptive history. Across five dexterous manipulation tasks, each stage addresses a complementary bottleneck: HOI-MMFP improves trajectory consistency and smoothness, task-space residuals provide the strongest grounding representation among the tested alternatives, and contact-dynamics adaptation enables robust real-world execution. Together, DexSynRefine improves real-world success rates over kinematic retargeting by 50-70~percentage points.
Hyesung Lee, Hyunwoo Jung, Si-Hwan Heo +1
Korea Institute of Science and Technology · KAIST · Hanyang University
High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand-object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task-relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video-to-dexterous-manipulation framework built around a shared interaction representation: stable object-side contacts recovered by aggregating noisy frame-wise observations in the canonical object space. These stable contacts serve a dual role: as trajectory-level constraints that guide reconstruction toward temporally coherent and physically plausible human HOI trajectories, and as explicit transfer targets for the dexterous hand, where Laplacian interaction optimization preserves the local hand-object geometry across embodiments and residual reinforcement learning refines the trajectory in simulation. Experiments on DexYCB and TACO show that C2Dex achieves end-to-end trajectory success rates of 57.78% and 26.67%, respectively, substantially outperforming the strongest baselines (17.78% and 10.00%) under identical evaluation criteria. Real-robot replay experiments further demonstrate physical feasibility across diverse contact-rich manipulation tasks. Project page: https://k-jie.github.io/C2Dex/
Jie Ren, Zhehao Jiang, Yinhong Yang +9
1Nanjing University, Nanjing, China. · 2China Mobile Research Institute, Beijing, China.