Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.
Figures & tables
Fig. 2: Overview of our framework. (a) Interaction-aware contact reconstruction combines hand–object segmentation and mesh-level contact prediction to recover temporally stable opposing contacts for robot grasp retargeting. (b) Depth-consistent compositing compares scene and rendered-robot depth to enforce physically correct robot–object occlusions. The resulting robotized videos are co-trained with real robot demonstrations for policy learning.
Policy
Training data
Adjust Bottle
Grab Roller
Place Burger Fries
Shake Bottle
Average
ID
OOD
ID
OOD
ID
OOD
ID
OOD
ID
OOD
DP-ViT
Baseline
33
15
27
1
40
21
69
29
42.25
16.50
Masquerade
76
37
85
12
37
23
64
22
65.50
23.50
Ours
82
70
75
33
46
43
69
37
68.00
45.75
DP-Swin
Baseline
73
30
80
22
39
10
84
10
69.00
18.00
Masquerade
94
57
94
43
69
39
84
36
85.25
43.75
TABLE I: Task success rates (%) on RoboTwin over 100 evaluation rollouts. All policies are trained in the ID scene. ID and OOD denote evaluation without and with visual distractors, respectively. Average values are computed across the four tasks. Bold indicates the best result for each policy backbone and evaluation setting.
Fig. 3: RoboTwin simulation evaluation scenes for Adjust Bottle, Grab Roller, Place Burger Fries, and Shake Bottle (top to bottom). The left column shows the ID scenes containing only task-relevant objects, while the right column shows the OOD scenes with additional visual distractors.
Variant
Adjust Bottle
Grab Roller
Place Burger Fries
Shake Bottle
Average
ID
OOD
ID
OOD
ID
OOD
ID
OOD
ID
OOD
Ours
82
70
75
33
46
43
69
37
68.00
45.75
w/o Contact Reconstruction
83
41
84
24
47
26
69
30
70.75
30.25
w/o Depth-Based Occlusion
83
65
83
13
46
36
53
29
66.25
35.75
TABLE II: Ablation results on RoboTwin using DP-ViT. All variants are trained in the ID scene. ID and OOD denote evaluation without and with visual distractors, respectively. Success rates (%) are measured over 100 rollouts, and averages are computed across the four tasks.
Fig. 4: Real-world evaluation scenes for Pick Up Sponge (top) and Pick Up Box (bottom). The left column shows the ID scenes containing only task-relevant objects, while the right column shows the OOD scenes with additional visual distractors.
Training data
Pick Up Box
Pick Up Sponge
ID
OOD
ID
OOD
Baseline
4 /10
1/10
8/10
5/10
Masquerade
3/10
2 /10
9 /10
7/10
Ours
4 /10
2 /10
9 /10
8 /10
TABLE III: Real-world task success over 10 trials per task and evaluation scene. ID and OOD denote evaluation without and with visual distractors, respectively. Bold indicates the best result in each setting.
Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning. Project webpage: https://roboedit.github.io/
Yaowei Guo, Zeng Tao, Yuxin Jiang +7
University of California, Los Angeles · Massachusetts Institute of Technology · University of Utah
How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present DO AS I DO, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. DO AS I DO reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, DO AS I DO outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show in experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation.
Bhawna Paliwal, Haritheja Etukuru, William Liang +3
Learning from human video demonstrations remains challenging due to noisy hand-object interactions, unseen objects with partial observation, and cross-embodiment discrepancy. To address these challenges, we present \textit{HOWTransfer} (\emph{H}and-\emph{O}bject \emph{O}pen-\emph{W}orld Transfer), a hand-centric framework that distills human demonstrations into contact-aware, taxonomy-informed, and diverse robotic trajectories. Instead of relying on object-specific descriptions, vision-language queries, or explicit object-state tracking, \emph{HOWTransfer} recovers temporally consistent 3D hand motion and localizes temporal contact intervals by reasoning over observed hand-object interaction cues. The localized contact onsets are then used to retarget human grasp intent into multi-modal parallel-jaw grasp hypotheses, which are propagated along the recovered wrist trajectory to generate robot-executable motions. Finally, a trajectory editing stage refines contact alignment and produces diverse executable variants from a single demonstration. Experiments across diverse manipulation tasks show that \emph{HOWTransfer} enables accurate contact localization and high-quality robot motion retargeting with 86% success, which is preferred over teleoperated trajectories in a blinded preference study.
Yitian Shi, Di Wen, Zhengqi Han +6
Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany