Organizations: The Hong Kong University of Science and Technology, Hong Kong SAR, China · Zhejiang University, Hangzhou, China · Tencent Robotics X, Shenzhen, China · Hunan University, Changsha, China
Humanoid whole-body teleoperation translates human motion into stable robot behavior in real time. Existing systems typically rely on online motion retargeting to bridge human--robot morphological differences, but this process adds latency and can produce physically infeasible targets. Meanwhile, diverse, noisy, and partial human-motion observations often fall outside the training distribution, potentially causing unstable robot behavior. We propose a retargeting-free policy that maps raw human motion directly to robot joint commands in a single forward pass, eliminating online kinematic adaptation. To improve robustness, we learn a codebook of full-body motion primitives that projects out-of-distribution observations onto plausible motion prototypes and recovers full-body motion from partial inputs. Experiments on a Unitree~G1 in simulation and on hardware, using virtual reality, optical mocap, text-to-motion generation, and monocular video inputs, show that our method outperforms baselines in latency and robustness.
Figures & tables
Fig. 1 : Success rate–latency tradeoff of different methods. Latency is the real-robot response delay, estimated by aligning operator and robot optical-flow trajectories. Our method, codebook , achieves the highest SR with the lowest latency among the compared methods.
Fig. 2 : Overview of the proposed teacher-student framework. Stage 1 trains a privileged MoE teacher with adaptive sampling; Stage 2 distills it into a codebook student through encoder–estimator codebook matching; at deployment, the estimator–decoder path maps operator observations and proprioceptive history to humanoid commands.
Method
SR ↑
MPJPE ↓
MPJVE ↓
MPKPE ↓
Eroot_pos↓
Eroot_lin_vel↓
Eroot_yaw↓
Eroot_yaw_vel↓
TWIST [ 37 ]
0.8451
0.2950
2.6869
0.2187
1.0449
0.7391
0.9395
1.3754
TWIST2 [ 38 ]
0.6199
0.3314
3.9878
0.3300
0.9958
1.0133
0.8999
1.9379
GMT [ 3 ]
0.7540
0.3122
2.8031
0.2453
0.9205
0.7825
0.8262
1.3557
CLONE [ 15 ]
0.8335
0.4636
3.3079
0.3358
0.7970
0.8804
0.6789
1.5397
ANY2TRACK [ 40 ]
0.0556
0.7279
7.6275
0.7699
1.5329
2.3334
2.3479
4.5856
MLP (ours, w/o codebook)
0.9529
0.2996
1.4412
0.2060
0.6016
0.4435
0.3761
1.1816
TABLE I : Task completion and penalized tracking errors on the SONIC bones-seed set, evaluated under a mild σ=0.1 m Gaussian keypoint perturbation applied uniformly to every method.
Method
Look-ahead
Retarget
Policy
E2E
TWIST [ 37 ]
0.0
21.930
0.442
22.372
TWIST2 [ 38 ]
0.0
22.615
0.731
23.346
GMT [ 3 ]
1900.0
22.170
0.880
1923.050
CLONE [ 15 ]
0.0
0.0
0.933
0.933
ANY2TRACK [ 40 ]
0.0
21.966
0.526
22.492
SONIC (G1) [ 18 ]
180.0
21.976
2.101
204.077
TABLE II : End-to-end online command latency (ms) under a shared simulated runtime.
Method
Clean ↑
Obs. dropout ↑
Keypoint noise ↑
Root-height drift ↑
Init. bias ↑
Unseen motion ↑
TWIST [ 37 ]
0.9114
0.9011
0.6344
0.8286
0.8942
0.4467
TWIST2 [ 38 ]
0.8482
0.8139
0.4090
0.0463
0.7895
0.4300
GMT [ 3 ]
0.8870
0.8838
0.4679
0.3770
0.8596
0.3967
CLONE [ 15 ]
0.8725
0.8401
0.6167
0.0000
0.8430
0.4700
ANY2TRACK [ 40 ]
0.3425
0.2913
0.0266
0.0536
0.3348
0.1500
MLP (ours, w/o codebook)
0.9584
0.9163
0.8675
0.8306
0.9535
0.5433
TABLE III : OOD robustness: SR at a representative scale of each condition—Obs. dropout s=0.50 , Keypoint noise s=0.30 m, Root-height drift s=0.20 m, Init. bias s=0.30 , Unseen motion is a single Motion-X split.
Fig. 3 : SR vs. keypoint Gaussian noise σ . The codebook student remains robust beyond σ=0.30 m, where baselines and the MLP variant degrade substantially.
Fig. 4 : Per- (c,k) code usage on the evaluation set for the clean encoder input (left) and partial, noisy estimator input (right). Rows are per-channel-normalized and the C=128 channels are sorted by the encoder’s argmax.
Fig. 5 : VR-headset end-effector tracking.
Fig. 6 : Full-body optical-mocap teleoperation.
Fig. 7 : Representative text-driven motion tracking results on the real Unitree G1.
TABLE IV : Teacher-side ablation on the cross-simulator held-out split. A.S. denotes adaptive sampling. Best results in bold .
C
D
C⋅D
SR ↑
MPJPE ↓
Top-1(%) ↓
Dead(%) ↓
64
6
384
0.939
0.296
56.9
1.0
64
12
768
0.932
0.291
46.7
1.8
128
12
1536
0.943
0.280
47.8
2.7
256
6
1536
0.936
0.286
60.4
1.4
128
24
3072
0.940
0.286
42.4
5.7
TABLE V : Codebook-capacity ablation on the SONIC bones-seed set. Top-1 is the largest per-channel code-use frequency, and Dead is the percentage of unused code slots. The selected configuration is shaded.
Fig. 9 : Effect of the number of input keypoints k . Each curve is normalized by its value at k=3 ; arrows indicate the preferred direction.
School of Electronic Science and Engineering, Nanjing University, Nanjing, China · Jiangsu Mobile Information System Integration Co., Ltd., Nanjing, China · China Mobile Zijin (Jiangsu) Innovation Research Institute Co., Ltd., Nanjing, China +1
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China