Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning. Project webpage: https://roboedit.github.io/
Figures & tables
Figure 1: RoboEdit Overview: The framework takes an RGB human video and a target robot, to generate a physically plausible robot interaction video and 3D robot-hand states (RoboEdit-Trans). These states enable motion supervision for downstream robot control, while the RoboEdit-ADC pipeline generates the large-scale paired dataset (RoboEdit-14M) required for training.
Figure 2: RoboEdit-14M spans diverse everyday manipulation tasks across a wide range of robot embodiments.
Dataset
Frames
View
Camera
Scene
Emb.
RGB Pair
Robot State
Auto
H&R ( Xie et al., 2025 )
∼ 1.2M
Third
Static
Real
1
✓
✓
×
H2R ( Li et al., 2026 )
∼ 1M
Ego
Moving
Real
3
✓
×
✓
UniDex ( Zhang et al., 2026 )
9M
Ego
Moving
Real
8
×
✓
×
X-Humanoid ( Yang et al., 2025 )
2.8M
Third
Moving
Synth.
1
✓
×
×
RoboEdit-14M (Ours)
14.1M
Ego+Third
Both
Both
7
✓
✓
✓
Table 1: Comparison of representative human-to-robot manipulation datasets. Frames follow each work’s reported scale; Emb. denotes the number of robot embodiments; Auto denotes curation without per-sample human intervention.
Figure 3: Qualitative comparison across four human-object interactions and four robot embodiments. The left two columns show RoboEdit-ADC results, while the remaining columns compare RoboEdit-Trans with baseline video editing models.
Figure 4: 3D robot-state predictions on RoboEdit-Trans edited videos across seven embodiments.
Method
Params
Cond.
Recon.
Local edit
VBench
OpenVE
SSIM ↑
LPIPS ↓
Edit LPIPS ↓
BG SSIM ↑
AQ ↑
DD ↑
MS ↑
Overall ↑
VACE
1.3B
S
0.8764
0.1070
0.0497
0.9487
0.4673
0.4867
0.9952
3.2270
UniVideo
14B
S
0.3249
0.6515
0.1043
0.4026
0.3871
0.4133
0.9890
2.2455
VINO
13B
S
0.5564
0.3348
0.0679
0.6041
0.4488
0.4433
0.9967
3.2711
Kiwi-Edit
5B
S
0.7415
0.2059
0.0627
0.8137
0.4307
0.4600
0.9950
3.2386
OmniWeaving
13B
S
0.8107
0.1590
0.0625
0.8855
0.4478
0.6900
0.9955
3.1634
Table 2: Quantitative results on RoboEdit-Trans benchmark. AQ/DD/MS denote VBench Aesthetic Quality/Dynamic Degree/Motion Smoothness; OpenVE is the overall OpenVE-Bench score. S/M/T: single reference / multi-keyframe / text instruction. (Best results: bold. Second-best: underlined.)
Figure 7Table 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Source dataset
Paired clips
Paired frames
Duration (h)
DexYCB
67,200
5,443,200
50.40
GigaHands
11,715
948,915
8.79
H2O
7,955
644,355
5.97
HOT3D
35,591
2,882,871
26.69
TACO
52,086
4,218,966
39.06
Total
174,547
14,138,307
130.91
Appendix
Table S1: RoboEdit-14M source distribution, with duration computed at 30 FPS.
Embodiment
Real
Synthetic
Total
Ability
24,197
4,839
29,036
Allegro
24,197
4,839
29,036
Unitree Dex3
12,237
2,447
14,684
Inspire
24,197
4,839
29,036
Panda gripper
12,237
2,447
14,684
SCHUNK SVH
24,197
4,838
29,035
Appendix
Table S2: RoboEdit-14M distribution by target embodiment.
Figure S1: Synthetic paired-video examples. Each column shows an aligned synthetic human frame and its robot target under the same scene, object state, and camera view.
Figure S2: Additional 3D Robot-State Decoder results on RoboEdit-Trans edited videos. Red overlays denote predicted camera-space hand states.
Method
Stage
Time (s)
RoboEdit-ADC
Object Segment
48.17
Object Pose
59.25
Object Mesh
125.58
Hand Pose
251.92
Camera State
104.43
Kinematic Retargeting
111.50
Appendix
Table S3: Runtime breakdown for RoboEdit-ADC and RoboEdit-Trans on one H100 GPU.
Figure S3: Additional qualitative comparisons on five human-object interactions.
Figure S4: Simulation rollouts of the trajectory-conditioned controller across four YCB-object manipulation tasks.
Figure S5: Additional real-robot deployment results across four YCB-object tasks.
Figure S6: Distribution of the 300-case benchmark. Legends show case counts and percentages; task families are coarse, source-based activity groups.
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodiment-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI.
Zhenjie Yang, Xingyu Jiao, Guopeng Zhong +18
Inspire Robots · Shanghai Jiao Tong University · Fudan University +3
Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled representations, where task-relevant information is coupled with human-specific kinematics, limiting their adaptability. We propose a generative framework for cross-embodiment video editing that directly addresses this by learning explicitly disentangled task and embodiment representations. Our method factorizes a demonstration video into two orthogonal latent spaces by enforcing a dual contrastive objective: it minimizes mutual information between the spaces to ensure independence while maximizing intra-space consistency to create stable representations. A parameter-efficient adapter injects these latent codes into a frozen video diffusion model, enabling the synthesis of a coherent robot execution video from a single human demonstration, without requiring paired cross-embodiment data. Experiments show our approach generates temporally consistent and morphologically accurate robot demonstrations, offering a scalable solution to leverage internet-scale human video for robot learning.
Zhiyuan Li, Wenyan Yang, Wenshuai Zhao +4
Department of Electrical Engineering and Automation, Aalto University, Finland · Department of Computer Science, Aalto University, Finland · Hong Kong University of Science and Technology, China +1
Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.