BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video
Authors: Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, +1 more
Organizations: School of Electronic Science and Engineering, Nanjing University, Nanjing, China · Jiangsu Mobile Information System Integration Co., Ltd., Nanjing, China · China Mobile Zijin (Jiangsu) Innovation Research Institute Co., Ltd., Nanjing, China · School of Intelligence Science and Technology, Nanjing University, Suzhou, China
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.
Figures & tables
Figure 2: Model architecture. Input RGB video is encoded by the visual and temporal encoders, and the resulting shared motion representation is directly decoded into the target robot motion. Contact-aware refinement further improves support and temporal consistency.
Figure 3: Unified human and robot supervision. Human keypoints are adapted to robot morphology. Keypoint positions and directed segments supervise humanoid robot pose.
Figure 4: Contact-aware kinematic refinement. The robot’s lower limb are divided into contact or support states based on explicit contact representation, and optimized via corresponding post-processing to improve the physical plausibility.
Method
MAMPJPE ↓
RTE ↓
Jitter ↓
Accel ↓
Fail 65 ↓
Fail 100 ↓
FS ↓
Exec. SR ↑
Exec. MAMPJPE ↓
GT → NMR
33.26
837.62
1.71
1.16
3.99
1.39
1.30
346/362
40.32
GT → GMR
32.02
36.01
10.65
2.27
0.00
0.00
0.82
347/362
35.93
GVHMR → NMR
39.19
701.13
1.63
1.56
6.21
2.22
2.28
340/362
43.32
GVHMR → GMR
39.46
637.33
7.94
2.51
3.43
0.28
3.30
313/362
41.77
WHAM → NMR
43.15
849.36
4.87
3.07
9.36
4.17
6.64
342/362
46.56
WHAM → GMR
45.88
217.76
63.26
14.57
10.47
6.30
13.30
295/362
42.89
Table 1: Quantitative comparison on G1. The oracle rows at the top use SMPL GT to isolate errors from human motion reconstruction, and thus are excluded from the ranking.
Figure 5: Qualitative results.
Method
H1
T1
Tienkung
G1
R1
GR1-T1
GR2-V3
Atlas
GVHMR → GMR
77.79
70.09
103.23
39.46
–
–
–
–
WHAM → GMR
99.66
77.26
115.86
45.88
–
–
–
–
BeyondRetarget
51.32
51.25
67.99
26.36
27.81
43.78
49.06
72.05
Table 2: Multi-robot MAMPJPE evaluation. The “–” entries denote robots unsupported by GMR.
Variant
MAMPJPE ↓
RTE ↓
TDE ↓
EDE ↓
Jitter ↓
Accel ↓
FS ↓
Viol. ↓
w/o non-uniform alignment
36.08
89.84
5.17
15.77
1.60
1.58
1.89
4.39
w/o Ldir
29.62
94.77
6.95
21.78
1.45
1.59
2.08
7.00
w/o Ltemp
26.44
77.12
2.95
15.41
1.93
1.61
1.99
5.66
w/o refinement
25.73
91.33
2.95
14.31
2.17
1.74
4.51
22.18
Full model
26.36
80.74
2.95
15.29
1.42
1.51
1.71
6.17
Table 3: Component ablation results.
Figure 6: Real-robot evaluation and visual teleoperation.
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France · IRT Saint-Exupéry, Toulouse, France · Artificial and Natural Intelligence Toulouse Institute (ANITI), Toulouse, France