Organizations: The Hong Kong University of Science and Technology, Hong Kong SAR, China · Zhejiang University, Hangzhou, China · Tencent Robotics X, Shenzhen, China · Hunan University, Changsha, China
Humanoid whole-body teleoperation translates human motion into stable robot behavior in real time. Existing systems typically rely on online motion retargeting to bridge human--robot morphological differences, but this process adds latency and can produce physically infeasible targets. Meanwhile, diverse, noisy, and partial human-motion observations often fall outside the training distribution, potentially causing unstable robot behavior. We propose a retargeting-free policy that maps raw human motion directly to robot joint commands in a single forward pass, eliminating online kinematic adaptation. To improve robustness, we learn a codebook of full-body motion primitives that projects out-of-distribution observations onto plausible motion prototypes and recovers full-body motion from partial inputs. Experiments on a Unitree~G1 in simulation and on hardware, using virtual reality, optical mocap, text-to-motion generation, and monocular video inputs, show that our method outperforms baselines in latency and robustness.
Figures & tables
Fig. 1 : Success rate–latency tradeoff of different methods. Latency is the real-robot response delay, estimated by aligning operator and robot optical-flow trajectories. Our method, codebook , achieves the highest SR with the lowest latency among the compared methods.
Fig. 2 : Overview of the proposed teacher-student framework. Stage 1 trains a privileged MoE teacher with adaptive sampling; Stage 2 distills it into a codebook student through encoder–estimator codebook matching; at deployment, the estimator–decoder path maps operator observations and proprioceptive history to humanoid commands.
Method
SR ↑
MPJPE ↓
MPJVE ↓
MPKPE ↓
Eroot_pos↓
Eroot_lin_vel↓
Eroot_yaw↓
Eroot_yaw_vel↓
TWIST [ 37 ]
0.8451
0.2950
2.6869
0.2187
1.0449
0.7391
0.9395
1.3754
TWIST2 [ 38 ]
0.6199
0.3314
3.9878
0.3300
0.9958
1.0133
0.8999
1.9379
GMT [ 3 ]
0.7540
0.3122
2.8031
0.2453
0.9205
0.7825
0.8262
1.3557
CLONE [ 15 ]
0.8335
0.4636
3.3079
0.3358
0.7970
0.8804
0.6789
1.5397
ANY2TRACK [ 40 ]
0.0556
0.7279
7.6275
0.7699
1.5329
2.3334
2.3479
4.5856
MLP (ours, w/o codebook)
0.9529
0.2996
1.4412
0.2060
0.6016
0.4435
0.3761
1.1816
TABLE I : Task completion and penalized tracking errors on the SONIC bones-seed set, evaluated under a mild σ=0.1 m Gaussian keypoint perturbation applied uniformly to every method.
Method
Look-ahead
Retarget
Policy
E2E
TWIST [ 37 ]
0.0
21.930
0.442
22.372
TWIST2 [ 38 ]
0.0
22.615
0.731
23.346
GMT [ 3 ]
1900.0
22.170
0.880
1923.050
CLONE [ 15 ]
0.0
0.0
0.933
0.933
ANY2TRACK [ 40 ]
0.0
21.966
0.526
22.492
SONIC (G1) [ 18 ]
180.0
21.976
2.101
204.077
TABLE II : End-to-end online command latency (ms) under a shared simulated runtime.
Method
Clean ↑
Obs. dropout ↑
Keypoint noise ↑
Root-height drift ↑
Init. bias ↑
Unseen motion ↑
TWIST [ 37 ]
0.9114
0.9011
0.6344
0.8286
0.8942
0.4467
TWIST2 [ 38 ]
0.8482
0.8139
0.4090
0.0463
0.7895
0.4300
GMT [ 3 ]
0.8870
0.8838
0.4679
0.3770
0.8596
0.3967
CLONE [ 15 ]
0.8725
0.8401
0.6167
0.0000
0.8430
0.4700
ANY2TRACK [ 40 ]
0.3425
0.2913
0.0266
0.0536
0.3348
0.1500
MLP (ours, w/o codebook)
0.9584
0.9163
0.8675
0.8306
0.9535
0.5433
TABLE III : OOD robustness: SR at a representative scale of each condition—Obs. dropout s=0.50 , Keypoint noise s=0.30 m, Root-height drift s=0.20 m, Init. bias s=0.30 , Unseen motion is a single Motion-X split.
Fig. 3 : SR vs. keypoint Gaussian noise σ . The codebook student remains robust beyond σ=0.30 m, where baselines and the MLP variant degrade substantially.
Fig. 4 : Per- (c,k) code usage on the evaluation set for the clean encoder input (left) and partial, noisy estimator input (right). Rows are per-channel-normalized and the C=128 channels are sorted by the encoder’s argmax.
Fig. 5 : VR-headset end-effector tracking.
Fig. 6 : Full-body optical-mocap teleoperation.
Fig. 7 : Representative text-driven motion tracking results on the real Unitree G1.
TABLE IV : Teacher-side ablation on the cross-simulator held-out split. A.S. denotes adaptive sampling. Best results in bold .
C
D
C⋅D
SR ↑
MPJPE ↓
Top-1(%) ↓
Dead(%) ↓
64
6
384
0.939
0.296
56.9
1.0
64
12
768
0.932
0.291
46.7
1.8
128
12
1536
0.943
0.280
47.8
2.7
256
6
1536
0.936
0.286
60.4
1.4
128
24
3072
0.940
0.286
42.4
5.7
TABLE V : Codebook-capacity ablation on the SONIC bones-seed set. Top-1 is the largest per-channel code-use frequency, and Dead is the percentage of unused code slots. The selected configuration is shaded.
Fig. 9 : Effect of the number of input keypoints k . Each curve is normalized by its value at k=3 ; arrows indicate the preferred direction.
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.
Tianyu Xiong, Yi Lu, Jinrui Wang +6
School of Electronic Science and Engineering, Nanjing University, Nanjing, China · Jiangsu Mobile Information System Integration Co., Ltd., Nanjing, China · China Mobile Zijin (Jiangsu) Innovation Research Institute Co., Ltd., Nanjing, China +1
Retargeting human motion to humanoid robots is critical for teleoperation, imitation learning and human-robot interaction. However, it remains challenging because of substantial morphological discrepancies between humans and robots, including differences in skeletal topology, limb proportions and degrees of freedom, as well as the scarcity of paired motion data. This paper presents Human2Humanoid, an unsupervised motion retargeting framework that transfers human motions to humanoid robot behaviors with high fidelity. To bridge the domain gap under unpaired data, we adopt a CycleGAN-based architecture equipped with a skeleton-aware graph convolutional network to capture topology-dependent motion features. To address cross-domain scale mismatches, we introduce a morphology-invariant end-effector consistency loss that aligns normalized end-effector trajectories to preserve motion semantics across embodiments. To improve physical plausibility and reduce contact artifacts, we impose explicit physics-aware feasibility constraints to encourage reproduction of the contact patterns in the source motion. Experimental results show that the proposed method successfully retargets human motion to the Unitree G1 humanoid robot without paired data, and outperforms existing methods in both downstream controllability and physical feasibility.
Tianchen Huang, Feiyang Yuan, Junchi Gu +5
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
Direct transfer from human demonstration to learnable robot action is a crucial step towards scalable whole-body mobile manipulation. While human data scales better than mobile teleoperation, it requires overcoming significant embodiment gaps. Existing retargeting methods yield imprecise or inconsistent solutions, causing action multi-modality that prevents supervised policies from reliably converging. We present Whole-body-Aware Retargeting from human Pose (WARP), an offline pipeline that explicitly models embodiment differences to extract precise, unique whole-body actions. WARP leverages a closed-form Shoulder-Elbow-Wrist (SEW) geometric solver for exact end-effector tracking while preserving whole-body structural intent. Paired with lazy mobile-base control, it extracts accurate, consistent robot trajectories. Evaluations show WARP provides highly reliable data for open-loop real-world replay. To our knowledge, WARP is the first framework to achieve zero-shot whole-body mobile manipulation directly from offline human demonstrations, eliminating the need for human-in-the-loop teleoperation action data. More details on https://warp-retarget.github.io/
Zhenyang Chen, Chuizheng Kong, Chuye Zhang +4
Georgia Institute of Technology, Atlanta, Georgia 30332