Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution. Prior methods either require task-specific RL training, limiting scalability, or assume clean motion-capture trajectories and thus cannot operate directly on video. We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands. The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement. Each stage optimizes the corresponding representation directly, correcting errors before they propagate downstream or must be absorbed by a learned policy. Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods. On 60 monocular video clips, we achieve 53% success, compared with 28% for the best prior video-to-robot pipeline. Its trajectories also execute on a real bimanual robot across diverse tasks.
Figures & tables
Figure 1: From human video to robot execution. Top: monocular RGB videos of human hand-object interaction. Bottom: the trajectories OmniHOI produces from them, executed on a pair of XHands.
Figure 2: OmniHOI is built on several key designs. Adaptive sampling reduces the video clip to keyframes (left) . Hand pose in each keyframe is reconstructed, refined and retargeted with contact information preserved (middle) . Interpolating between the keyframe configurations then gives an initial robot trajectory. Physics-in-the-loop refinement then rolls the trajectory out in simulation and updates the keyframe configurations with CMA-ES, yielding a task-effective trajectory (right) .
Figure 3: Keyframes are equal increments of At . Equal steps ΔA are unequal in time: keyframes crowd where the hand moves fast, unlike uniform sampling of the same budget.
Figure 4: Occlusion-order loss. Input image and handle close-ups before (a) and after (b) refinement. A ray through an overlap pixel meets the hand at depth Dh and the object at depth DO ; the order is violated in (a) and correct in (b), where the fingers no longer intersect the handle.
Figure 5: Contact-aware retargeting. Left: Arrows denote the gradient directions provided by the SDF. Gradients acting on the same finger link point in opposite directions and cancel each other out, but the repair makes the finger be effectively pushed out of the object along the bone chain. Middle: contact guidance: the anchors are sampled on the robot surface within the band around the contact region Rf , which is required to reach. Right: comparisons between ground-truth MANO hand, the retargeting solver’s output, and the result after repair. With repair, the robot finger remains penetration-free even when the noisy ground-truth MANO annotation already penetrates the object.
Figure 6: Qualitative hand-object reconstruction from a single image. FoHo: FollowMyHold.
Overall
ARCTIC
Method
F5
F10
CD
MPVPE
MPJPE
IV
F5
F10
CD
MPVPE
MPJPE
IV
↑
↑
(cm 2 ) ↓
(mm) ↓
(mm) ↓
(cm 3 ) ↓
↑
↑
(cm 2 ) ↓
(mm) ↓
(mm) ↓
(cm 3 ) ↓
EasyHOI
0.2187
0.3890
1.7110
14.63
14.02
23.66
0.1720
0.3067
2.6391
16.20
15.43
27.19
FollowMyHold
0.2007
0.3602
1.7731
8.92
8.48
19.73
0.1698
0.3035
2.6325
10.69
10.10
28.30
iHOI
0.2158
0.3881
1.7384
11.41
10.88
2.90
0.1741
0.3137
2.6672
12.79
12.07
2.35
OmniHOI (ours)
0.3878
0.5831
0.9574
6.05
5.71
2.75
0.2823
0.4568
1.6048
5.36
4.99
2.64
Table 1: Reconstruction results. F5/F10: object F-score at a 5/10 mm surface-distance threshold; CD: object Chamfer distance; MPVPE/MPJPE: mean per-vertex/per-joint error of the hand with dataset ground truth after Procrustes alignment; IV: hand-object intersection volume. Deeper color is better.
Figure 7: The 150 transfer tasks by source dataset. (a) Action category, where “other” is cutting and measuring. (b) Duration of the replayed in-hand clips. (c) Single-hand versus bimanual tasks.
Inspire Hand (6DoF)
XHand (12DoF)
Wuji Hand (20DoF)
SharpaWave (22DoF)
Shadow Hand (20DoF)
Method
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Time
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(min) ↓
ManipTrans
24
9.4
37
22
9.7
33
–
–
–
–
–
–
26
10.7
34
385
SPIDER
26
6.6
43
26
6.2
43
20
8.3
43
19
8.1
54
15
8.5
45
37
OmniRetarget
24
8.4
33
31
3.8
35
27
6.4
43
20
7.3
46
16
8.1
50
0.3
OmniHOI (ours)
39
6.7
26
89
1.8
12
86
1.6
11
89
1.5
10
85
2.4
14
19
Table 2: Human-to-robot motion transfer on 150 trajectories per hand across different robot hands. ManipTrans does not support Wuji Hand and SharpaWave. Time is the mean wall-clock minutes to process a trajectory over all hands a method was run on, measured on RTX 4090. It covers per-task RL training and rollout for ManipTrans (excluding imitator pre-training), whole pipeline for SPIDER, and human-to-robot motion transfer for OmniHOI . Deeper color is better.
Method
Succ.
Succ.Prog.
Trans.
Rot.
(%) ↑
(%) ↑
(cm) ↓
( ∘ ) ↓
VideoManip
28
49
8.6
31
Do as I Do
25
48
9.1
31
OmniHOI (ours)
53
71
5.7
26
Table 3: End-to-end video-to-robot transfer on the 60 clips, under the ATE criterion of Section 4.2 . Succ.Prog.: Progress fraction of the longest prefix that passes the success criterion (mean).
Figure 8: Real-world deployment. Two representative tasks, sweeping and stacking, in temporal order from left to right. Top: the robot executing the trajectory. Bottom: input human video.
Variant
PA-MPVPE
PA-MPJPE
IV
(mm) ↓
(mm) ↓
(cm 3 ) ↓
Placement only
5.37
5.03
4.27
+ image alignment
5.84
5.50
3.61
+ contact-penetration
5.81
5.48
3.03
+ finger-crossing repair
6.05
5.71
2.75
Table 4: Reconstruction ablations on the 1500 frames of Table 1 : starting from the hand placed as in Section 3.1 , the refinement steps of Section 3.1 (image alignment with Lord , contact-penetration with Lgrasp , finger-crossing repair) are added one at a time.
Inspire Hand (6DoF)
XHand (12DoF)
Wuji Hand (20DoF)
SharpaWave (22DoF)
Shadow Hand (20DoF)
Variant
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
Succ.
Trans.
Rot.
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
Retargeting ( qinit )
25
8.8
34
29
4.6
38
22
6.8
40
17
7.0
45
12
8.5
53
+ contact refinement
21
8.3
34
23
5.4
44
19
7.5
46
17
7.7
50
9
8.5
59
+ physics-in-the-loop
36
6.8
27
87
1.7
12
88
1.5
10
87
1.8
12
85
2.4
14
+ both
39
6.7
26
89
1.8
12
86
1.6
11
89
1.5
10
85
2.4
14
Table 5: Transfer ablations: the two refinement steps of Section 3.2 switched on and off, on the protocol of Table 2 (150 tasks per hand, from ground-truth MANO, ATE criterion). Deeper color is better.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Segmentation
SAM 3.1; objects are grounded by their category words and propagated through the clip
Depth and camera
MoGe-3 (ViT-G); focal length is the median estimate over three frames, principal point at the image center
Frame alignment
similarity to the first keyframe on table pixels within 400 px of the objects, 20% residual trimming
Object mesh
SAM 3D on the frame with at most 5% occlusion and the largest visible mask, decimated to 40k faces
Object pose
FoundationPose from the object mask in every frame, 5 refinement iterations
Hand placement
3 rounds of the similarity fit with 10% residual trimming; κ=2.41 , i.e. wi=0.3 at di=ℓ/2
Appendix
Table 6: Key reconstruction settings.
Adaptive
Uniform
r
Succ.
Trans.
Rot.
Rollouts
Succ.
Trans.
Rot.
Rollouts
(%) ↑
(cm) ↓
( ∘ ) ↓
(%) ↑
(cm) ↓
( ∘ ) ↓
0.05
57
4.4
29
0.9K
50
5.6
33
1.0K
0.1
84
1.9
13
1.9K
82
2.3
14
1.7K
0.2
83
1.9
12
3.4K
81
1.9
13
3.4K
1
–
–
–
–
89
1.7
12
13K
Appendix
Table 7: Keyframe sampling on XHand, 100 tasks, ATE criterion. At each ratio r both rules select the same number of keyframes and keep the first and last frame. Rollouts: number of physics simulations, i.e., candidate actions the optimizer tries out in simulation, which measures the actual compute (median over tasks).
Component
Setting
Solver
timestep 1/240 s (4 steps per 60 Hz frame); implicit-fast integrator; Newton solver with 10 iterations and 10 no-slip iterations; elliptic friction cones, impedance ratio 10
Gravity
9.81 m/s 2 along the negative normal of the support plane
Contacts
condim 6; friction (1.0, 0.01, 0.0002); solref (0.003, 1), solimp (0.9, 0.99, 0.001); margin 1 mm; the hands do not collide with the table
Table
plane taken from the ground-truth scene
Objects
ground-truth initial pose and velocity; cataloged masses (typical weights for TACO and OakInk2, manufacturer weights for ARCTIC) spread uniformly over CoACD ( Wei et al., 2022 ) hulls (2 mm, at most 32); objects that neither move nor are touched in the ground truth are fixed; every annotated object is present
Articulations
ARCTIC objects: limited hinge over the recorded range, damping 0.02 N m s, no friction loss
Appendix
Table 8: Simulation settings shared by all methods in the transfer benchmark.
Figure 9: Real-world film strips : insertion, stacking, sweeping and scanning. For each task, top: the input human video; bottom: the robot executing the trajectory. Time runs from left to right.
Figure 10: Real-world film strips , continued: scrubbing and the two pouring tasks, laid out as in Figure 9 .