BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video
Authors: Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, +1 more
Organizations: School of Electronic Science and Engineering, Nanjing University, Nanjing, China · Jiangsu Mobile Information System Integration Co., Ltd., Nanjing, China · China Mobile Zijin (Jiangsu) Innovation Research Institute Co., Ltd., Nanjing, China · School of Intelligence Science and Technology, Nanjing University, Suzhou, China
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.
Figures & tables
Figure 2: Model architecture. Input RGB video is encoded by the visual and temporal encoders, and the resulting shared motion representation is directly decoded into the target robot motion. Contact-aware refinement further improves support and temporal consistency.
Figure 3: Unified human and robot supervision. Human keypoints are adapted to robot morphology. Keypoint positions and directed segments supervise humanoid robot pose.
Figure 4: Contact-aware kinematic refinement. The robot’s lower limb are divided into contact or support states based on explicit contact representation, and optimized via corresponding post-processing to improve the physical plausibility.
Method
MAMPJPE ↓
RTE ↓
Jitter ↓
Accel ↓
Fail 65 ↓
Fail 100 ↓
FS ↓
Exec. SR ↑
Exec. MAMPJPE ↓
GT → NMR
33.26
837.62
1.71
1.16
3.99
1.39
1.30
346/362
40.32
GT → GMR
32.02
36.01
10.65
2.27
0.00
0.00
0.82
347/362
35.93
GVHMR → NMR
39.19
701.13
1.63
1.56
6.21
2.22
2.28
340/362
43.32
GVHMR → GMR
39.46
637.33
7.94
2.51
3.43
0.28
3.30
313/362
41.77
WHAM → NMR
43.15
849.36
4.87
3.07
9.36
4.17
6.64
342/362
46.56
WHAM → GMR
45.88
217.76
63.26
14.57
10.47
6.30
13.30
295/362
42.89
Table 1: Quantitative comparison on G1. The oracle rows at the top use SMPL GT to isolate errors from human motion reconstruction, and thus are excluded from the ranking.
Figure 5: Qualitative results.
Method
H1
T1
Tienkung
G1
R1
GR1-T1
GR2-V3
Atlas
GVHMR → GMR
77.79
70.09
103.23
39.46
–
–
–
–
WHAM → GMR
99.66
77.26
115.86
45.88
–
–
–
–
BeyondRetarget
51.32
51.25
67.99
26.36
27.81
43.78
49.06
72.05
Table 2: Multi-robot MAMPJPE evaluation. The “–” entries denote robots unsupported by GMR.
Variant
MAMPJPE ↓
RTE ↓
TDE ↓
EDE ↓
Jitter ↓
Accel ↓
FS ↓
Viol. ↓
w/o non-uniform alignment
36.08
89.84
5.17
15.77
1.60
1.58
1.89
4.39
w/o Ldir
29.62
94.77
6.95
21.78
1.45
1.59
2.08
7.00
w/o Ltemp
26.44
77.12
2.95
15.41
1.93
1.61
1.99
5.66
w/o refinement
25.73
91.33
2.95
14.31
2.17
1.74
4.51
22.18
Full model
26.36
80.74
2.95
15.29
1.42
1.51
1.71
6.17
Table 3: Component ablation results.
Figure 6: Real-robot evaluation and visual teleoperation.
Retargeting human motion to humanoid robots is critical for teleoperation, imitation learning and human-robot interaction. However, it remains challenging because of substantial morphological discrepancies between humans and robots, including differences in skeletal topology, limb proportions and degrees of freedom, as well as the scarcity of paired motion data. This paper presents Human2Humanoid, an unsupervised motion retargeting framework that transfers human motions to humanoid robot behaviors with high fidelity. To bridge the domain gap under unpaired data, we adopt a CycleGAN-based architecture equipped with a skeleton-aware graph convolutional network to capture topology-dependent motion features. To address cross-domain scale mismatches, we introduce a morphology-invariant end-effector consistency loss that aligns normalized end-effector trajectories to preserve motion semantics across embodiments. To improve physical plausibility and reduce contact artifacts, we impose explicit physics-aware feasibility constraints to encourage reproduction of the contact patterns in the source motion. Experimental results show that the proposed method successfully retargets human motion to the Unitree G1 humanoid robot without paired data, and outperforms existing methods in both downstream controllability and physical feasibility.
Tianchen Huang, Feiyang Yuan, Junchi Gu +5
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches. Standard approaches rely on Geometric Retargeting or Indirect Dynamic Retargeting pipelines. We identify that these intermediate kinematic projections introduce a geometric bias, restricting the search space and yielding suboptimal dynamic behaviors. In this paper, we propose Direct Dynamic Retargeting (DDR), a novel single-stage framework that generates high-fidelity, dynamically feasible trajectories directly from expert videos. By formulating the problem in the task space and leveraging a sampling-based Model Predictive Control solver within a physics simulator, DDR natively optimizes over complex contact sequences while mitigating input drift. Our experiments demonstrate that bypassing the geometric bias allows DDR to outperform state-of-the-art baselines in demonstration tracking accuracy. Furthermore, we establish that providing such physically viable references to RL agents accelerates training convergence and enhances the final execution of agile and balancing behaviors. Source code will be made publicly available.
Constant Roux, Ludovic De Matteïs, Armand Jordana +4
LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France · IRT Saint-Exupéry, Toulouse, France · Artificial and Natural Intelligence Toulouse Institute (ANITI), Toulouse, France
Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer. Frame-wise body estimates, video-level observations, and detailed hand evidence jointly constrain a single differentiable Momentum Human Rig (MHR) state, while transient hand artifacts are repaired in parameter space. The reconstructed motion is represented by arm-segment directions, elbow configuration, relative palm orientation, and bilateral wrist relations, and is realized on the target robot through multi-stage inverse kinematics and robot-specific hand adaptation. Within the broader system, Across-VAM provides video generation, whereas Across-WAM performs human-to-robot motion mapping. The method is evaluated on 16 monocular videos comprising 1,769 source frames, including 10 signing and six reach-to-grasp sequences. Unified reconstruction reduces mean hand reprojection error from 22.36 to 7.21 pixels relative to SAM 3D Body. All 16 retargeted trajectories completed kinematic simulation playback, and representative signing and reach-to-grasp motions were further demonstrated on a physical robot. The results demonstrate a unified pipeline from monocular human video to coordinated upper-body motion on a dual-arm dexterous robot.
Xiaoyu Yang, Sen Han, Da Li +1
Across Physics, Beijing, China · Beijing SFlare Robotics Technology Co., Ltd., Beijing, China