Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution. Prior methods either require task-specific RL training, limiting scalability, or assume clean motion-capture trajectories and thus cannot operate directly on video. We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands. The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement. Each stage optimizes the corresponding representation directly, correcting errors before they propagate downstream or must be absorbed by a learned policy. Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods. On 60 monocular video clips, we achieve 53% success, compared with 28% for the best prior video-to-robot pipeline. Its trajectories also execute on a real bimanual robot across diverse tasks.
Figures & tables
Figure 1: From human video to robot execution. Top: monocular RGB videos of human hand-object interaction. Bottom: the trajectories OmniHOI produces from them, executed on a pair of XHands.
Figure 2: OmniHOI is built on several key designs. Adaptive sampling reduces the video clip to keyframes (left) . Hand pose in each keyframe is reconstructed, refined and retargeted with contact information preserved (middle) . Interpolating between the keyframe configurations then gives an initial robot trajectory. Physics-in-the-loop refinement then rolls the trajectory out in simulation and updates the keyframe configurations with CMA-ES, yielding a task-effective trajectory (right) .
Figure 3: Keyframes are equal increments of At . Equal steps ΔA are unequal in time: keyframes crowd where the hand moves fast, unlike uniform sampling of the same budget.
Figure 4: Occlusion-order loss. Input image and handle close-ups before (a) and after (b) refinement. A ray through an overlap pixel meets the hand at depth Dh and the object at depth DO ; the order is violated in (a) and correct in (b), where the fingers no longer intersect the handle.
Figure 5: Contact-aware retargeting. Left: Arrows denote the gradient directions provided by the SDF. Gradients acting on the same finger link point in opposite directions and cancel each other out, but the repair makes the finger be effectively pushed out of the object along the bone chain. Middle: contact guidance: the anchors are sampled on the robot surface within the band around the contact region Rf , which is required to reach. Right: comparisons between ground-truth MANO hand, the retargeting solver’s output, and the result after repair. With repair, the robot finger remains penetration-free even when the noisy ground-truth MANO annotation already penetrates the object.
Figure 6: Qualitative hand-object reconstruction from a single image. FoHo: FollowMyHold.
Overall
ARCTIC
Method
F5
F10
CD
MPVPE
MPJPE
IV
F5
F10
CD
MPVPE
MPJPE
IV
↑
↑
(cm 2 ) ↓
(mm) ↓
(mm) ↓
(cm 3 ) ↓
↑
↑
(cm 2 ) ↓
(mm) ↓
(mm) ↓
(cm 3 ) ↓
EasyHOI
0.2187
0.3890
1.7110
14.63
14.02
23.66
0.1720
0.3067
2.6391
16.20
15.43
27.19
FollowMyHold
0.2007
0.3602
1.7731
8.92
8.48
19.73
0.1698
0.3035
2.6325
10.69
10.10
28.30
iHOI
0.2158
0.3881
1.7384
11.41
10.88
2.90
0.1741
0.3137
2.6672
12.79
12.07
2.35
OmniHOI (ours)
0.3878
0.5831
0.9574
6.05
5.71
2.75
0.2823
0.4568
1.6048
5.36
4.99
2.64
Table 1: Reconstruction results. F5/F10: object F-score at a 5/10 mm surface-distance threshold; CD: object Chamfer distance; MPVPE/MPJPE: mean per-vertex/per-joint error of the hand with dataset ground truth after Procrustes alignment; IV: hand-object intersection volume. Deeper color is better.
Figure 7: The 150 transfer tasks by source dataset. (a) Action category, where “other” is cutting and measuring. (b) Duration of the replayed in-hand clips. (c) Single-hand versus bimanual tasks.
Inspire Hand (6DoF)
XHand (12DoF)
Wuji Hand (20DoF)
SharpaWave (22DoF)
Shadow Hand (20DoF)
Method
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Time
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(min) ↓
ManipTrans
24
9.4
37
22
9.7
33
–
–
–
–
–
–
26
10.7
34
385
SPIDER
26
6.6
43
26
6.2
43
20
8.3
43
19
8.1
54
15
8.5
45
37
OmniRetarget
24
8.4
33
31
3.8
35
27
6.4
43
20
7.3
46
16
8.1
50
0.3
OmniHOI (ours)
39
6.7
26
89
1.8
12
86
1.6
11
89
1.5
10
85
2.4
14
19
Table 2: Human-to-robot motion transfer on 150 trajectories per hand across different robot hands. ManipTrans does not support Wuji Hand and SharpaWave. Time is the mean wall-clock minutes to process a trajectory over all hands a method was run on, measured on RTX 4090. It covers per-task RL training and rollout for ManipTrans (excluding imitator pre-training), whole pipeline for SPIDER, and human-to-robot motion transfer for OmniHOI . Deeper color is better.
Method
Succ.
Succ.Prog.
Trans.
Rot.
(%) ↑
(%) ↑
(cm) ↓
( ∘ ) ↓
VideoManip
28
49
8.6
31
Do as I Do
25
48
9.1
31
OmniHOI (ours)
53
71
5.7
26
Table 3: End-to-end video-to-robot transfer on the 60 clips, under the ATE criterion of Section 4.2 . Succ.Prog.: Progress fraction of the longest prefix that passes the success criterion (mean).
Figure 8: Real-world deployment. Two representative tasks, sweeping and stacking, in temporal order from left to right. Top: the robot executing the trajectory. Bottom: input human video.
Variant
PA-MPVPE
PA-MPJPE
IV
(mm) ↓
(mm) ↓
(cm 3 ) ↓
Placement only
5.37
5.03
4.27
+ image alignment
5.84
5.50
3.61
+ contact-penetration
5.81
5.48
3.03
+ finger-crossing repair
6.05
5.71
2.75
Table 4: Reconstruction ablations on the 1500 frames of Table 1 : starting from the hand placed as in Section 3.1 , the refinement steps of Section 3.1 (image alignment with Lord , contact-penetration with Lgrasp , finger-crossing repair) are added one at a time.
Inspire Hand (6DoF)
XHand (12DoF)
Wuji Hand (20DoF)
SharpaWave (22DoF)
Shadow Hand (20DoF)
Variant
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
Retargeting ( qinit )
25
8.8
34
29
4.6
38
22
6.8
40
17
7.0
45
12
8.5
53
+ contact refinement
21
8.3
34
23
5.4
44
19
7.5
46
17
7.7
50
9
8.5
59
+ physics-in-the-loop
36
6.8
27
87
1.7
12
88
1.5
10
87
1.8
12
85
2.4
14
+ both
39
6.7
26
89
1.8
12
86
1.6
11
89
1.5
10
85
2.4
14
Table 5: Transfer ablations: the two refinement steps of Section 3.2 switched on and off, on the protocol of Table 2 (150 tasks per hand, from ground-truth MANO, ATE criterion). Deeper color is better.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Segmentation
SAM 3.1; objects are grounded by their category words and propagated through the clip
Depth and camera
MoGe-3 (ViT-G); focal length is the median estimate over three frames, principal point at the image center
Frame alignment
similarity to the first keyframe on table pixels within 400 px of the objects, 20% residual trimming
Object mesh
SAM 3D on the frame with at most 5% occlusion and the largest visible mask, decimated to 40k faces
Object pose
FoundationPose from the object mask in every frame, 5 refinement iterations
Hand placement
3 rounds of the similarity fit with 10% residual trimming; κ=2.41 , i.e. wi=0.3 at di=ℓ/2
Appendix
Table 6: Key reconstruction settings.
Adaptive
Uniform
r
Succ.
Trans.
Rot.
Rollouts
Succ.
Trans.
Rot.
Rollouts
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
0.05
57
4.4
29
0.9K
50
5.6
33
1.0K
0.1
84
1.9
13
1.9K
82
2.3
14
1.7K
0.2
83
1.9
12
3.4K
81
1.9
13
3.4K
1
–
–
–
–
89
1.7
12
13K
Appendix
Table 7: Keyframe sampling on XHand, 100 tasks, ATE criterion. At each ratio r both rules select the same number of keyframes and keep the first and last frame. Rollouts: number of physics simulations, i.e., candidate actions the optimizer tries out in simulation, which measures the actual compute (median over tasks).
Component
Setting
Solver
timestep 1/240 s (4 steps per 60 Hz frame); implicit-fast integrator; Newton solver with 10 iterations and 10 no-slip iterations; elliptic friction cones, impedance ratio 10
Gravity
9.81 m/s 2 along the negative normal of the support plane
Contacts
condim 6; friction (1.0, 0.01, 0.0002); solref (0.003, 1), solimp (0.9, 0.99, 0.001); margin 1 mm; the hands do not collide with the table
Table
plane taken from the ground-truth scene
Objects
ground-truth initial pose and velocity; cataloged masses (typical weights for TACO and OakInk2, manufacturer weights for ARCTIC) spread uniformly over CoACD ( Wei et al., 2022 ) hulls (2 mm, at most 32); objects that neither move nor are touched in the ground truth are fixed; every annotated object is present
Articulations
ARCTIC objects: limited hinge over the recorded range, damping 0.02 N m s, no friction loss
Appendix
Table 8: Simulation settings shared by all methods in the transfer benchmark.
Figure 9: Real-world film strips : insertion, stacking, sweeping and scanning. For each task, top: the input human video; bottom: the robot executing the trajectory. Time runs from left to right.
Figure 10: Real-world film strips , continued: scrubbing and the two pouring tasks, laid out as in Figure 9 .
We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamics derive hand generalized forces and object wrenches for dynamics-level supervision. We further couple hand and object dynamics through contact-force transfer and recover active hand actuation as an interaction-level physical quantity. We formulate its empirical magnitude distribution into a probabilistic prior that penalizes unlikely actuation and suppresses mechanically implausible reconstructed motion. Experiments on three HOI datasets show consistent improvements in both hand and object reconstruction. The reconstructed trajectories further benefit downstream applications including hand world-model generation and robotic manipulation learning, demonstrating the value of physics-aware HOI modeling beyond reconstruction.
High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand-object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task-relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video-to-dexterous-manipulation framework built around a shared interaction representation: stable object-side contacts recovered by aggregating noisy frame-wise observations in the canonical object space. These stable contacts serve a dual role: as trajectory-level constraints that guide reconstruction toward temporally coherent and physically plausible human HOI trajectories, and as explicit transfer targets for the dexterous hand, where Laplacian interaction optimization preserves the local hand-object geometry across embodiments and residual reinforcement learning refines the trajectory in simulation. Experiments on DexYCB and TACO show that C2Dex achieves end-to-end trajectory success rates of 57.78% and 26.67%, respectively, substantially outperforming the strongest baselines (17.78% and 10.00%) under identical evaluation criteria. Real-robot replay experiments further demonstrate physical feasibility across diverse contact-rich manipulation tasks. Project page: https://k-jie.github.io/C2Dex/
Jie Ren, Zhehao Jiang, Yinhong Yang +9
1Nanjing University, Nanjing, China. · 2China Mobile Research Institute, Beijing, China.
How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present DO AS I DO, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. DO AS I DO reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, DO AS I DO outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show in experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation.
Bhawna Paliwal, Haritheja Etukuru, William Liang +3