Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning. Project webpage: https://roboedit.github.io/
Figures & tables
Figure 1: RoboEdit Overview: The framework takes an RGB human video and a target robot, to generate a physically plausible robot interaction video and 3D robot-hand states (RoboEdit-Trans). These states enable motion supervision for downstream robot control, while the RoboEdit-ADC pipeline generates the large-scale paired dataset (RoboEdit-14M) required for training.
Figure 2: RoboEdit-14M spans diverse everyday manipulation tasks across a wide range of robot embodiments.
Dataset
Frames
View
Camera
Scene
Emb.
RGB Pair
Robot State
Auto
H&R ( Xie et al., 2025 )
∼ 1.2M
Third
Static
Real
1
✓
✓
×
H2R ( Li et al., 2026 )
∼ 1M
Ego
Moving
Real
3
✓
×
✓
UniDex ( Zhang et al., 2026 )
9M
Ego
Moving
Real
8
×
✓
×
X-Humanoid ( Yang et al., 2025 )
2.8M
Third
Moving
Synth.
1
✓
×
×
RoboEdit-14M (Ours)
14.1M
Ego+Third
Both
Both
7
✓
✓
✓
Table 1: Comparison of representative human-to-robot manipulation datasets. Frames follow each work’s reported scale; Emb. denotes the number of robot embodiments; Auto denotes curation without per-sample human intervention.
Figure 3: Qualitative comparison across four human-object interactions and four robot embodiments. The left two columns show RoboEdit-ADC results, while the remaining columns compare RoboEdit-Trans with baseline video editing models.
Figure 4: 3D robot-state predictions on RoboEdit-Trans edited videos across seven embodiments.
Method
Params
Cond.
Recon.
Local edit
VBench
OpenVE
SSIM ↑
LPIPS ↓
Edit LPIPS ↓
BG SSIM ↑
AQ ↑
DD ↑
MS ↑
Overall ↑
VACE
1.3B
S
0.8764
0.1070
0.0497
0.9487
0.4673
0.4867
0.9952
3.2270
UniVideo
14B
S
0.3249
0.6515
0.1043
0.4026
0.3871
0.4133
0.9890
2.2455
VINO
13B
S
0.5564
0.3348
0.0679
0.6041
0.4488
0.4433
0.9967
3.2711
Kiwi-Edit
5B
S
0.7415
0.2059
0.0627
0.8137
0.4307
0.4600
0.9950
3.2386
OmniWeaving
13B
S
0.8107
0.1590
0.0625
0.8855
0.4478
0.6900
0.9955
3.1634
Table 2: Quantitative results on RoboEdit-Trans benchmark. AQ/DD/MS denote VBench Aesthetic Quality/Dynamic Degree/Motion Smoothness; OpenVE is the overall OpenVE-Bench score. S/M/T: single reference / multi-keyframe / text instruction. (Best results: bold. Second-best: underlined.)
Figure 7Table 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Source dataset
Paired clips
Paired frames
Duration (h)
DexYCB
67,200
5,443,200
50.40
GigaHands
11,715
948,915
8.79
H2O
7,955
644,355
5.97
HOT3D
35,591
2,882,871
26.69
TACO
52,086
4,218,966
39.06
Total
174,547
14,138,307
130.91
Appendix
Table S1: RoboEdit-14M source distribution, with duration computed at 30 FPS.
Embodiment
Real
Synthetic
Total
Ability
24,197
4,839
29,036
Allegro
24,197
4,839
29,036
Unitree Dex3
12,237
2,447
14,684
Inspire
24,197
4,839
29,036
Panda gripper
12,237
2,447
14,684
SCHUNK SVH
24,197
4,838
29,035
Appendix
Table S2: RoboEdit-14M distribution by target embodiment.
Figure S1: Synthetic paired-video examples. Each column shows an aligned synthetic human frame and its robot target under the same scene, object state, and camera view.
Figure S2: Additional 3D Robot-State Decoder results on RoboEdit-Trans edited videos. Red overlays denote predicted camera-space hand states.
Method
Stage
Time (s)
RoboEdit-ADC
Object Segment
48.17
Object Pose
59.25
Object Mesh
125.58
Hand Pose
251.92
Camera State
104.43
Kinematic Retargeting
111.50
Appendix
Table S3: Runtime breakdown for RoboEdit-ADC and RoboEdit-Trans on one H100 GPU.
Figure S3: Additional qualitative comparisons on five human-object interactions.
Figure S4: Simulation rollouts of the trajectory-conditioned controller across four YCB-object manipulation tasks.
Figure S5: Additional real-robot deployment results across four YCB-object tasks.
Figure S6: Distribution of the 300-case benchmark. Legends show case counts and percentages; task families are coarse, source-based activity groups.
Department of Electrical Engineering and Automation, Aalto University, Finland · Department of Computer Science, Aalto University, Finland · Hong Kong University of Science and Technology, China +1