Organizations: The Institute of Artificial Intelligence, China Telecom (TeleAI) · Shanghai Jiao Tong University · The University of Hong Kong · Zhejiang University · γ-Robotics
Whole-body teleoperation requires a humanoid robot to reproduce a human operator's behavior even when their terrains differ. This demands that the robot perceive local terrain and adapt its posture and contacts accordingly, rather than copy the operator's motion frame by frame. However, paired motion data linking the same behaviors across flat ground and different terrains remain scarce, limiting supervision for learning terrain-adaptive control. To enable whole-body teleoperation across mismatched terrains, we introduce NEXUS, a perceptive whole-body control framework that combines human motion commands with onboard sensory feedback. We first develop a scalable terrain-aware adaptation algorithm that efficiently generates high-quality motion pairs across motions and terrains without per-motion or per-terrain tuning. Using a paired motion corpus totaling nearly 1,000 hours, we train a perceptive whole-body controller through teacher-student learning to reproduce commanded behaviors across terrains. Experiments demonstrate efficient, scalable generation of high-quality motion data and show that NEXUS combines broad behavioral coverage with terrain adaptability and tracking fidelity, outperforming existing whole-body controllers on the evaluated benchmarks. Zero-shot real-world deployment enables real-time whole-body teleoperation on diverse unseen terrains, further validating the generalization of our method. Project website: https://nexus-humanoid.github.io/
Figures & tables
Fig. 2: NEXUS overview. (a) Contact-guided motion adaptation: contact detection, pelvis adjustment, IK with penetration correction, and trajectory reconstruction from contact keyframes. Contact markers are illustrative. (b) A privileged teacher trained with PPO supervises a student through DAgger-style behavior cloning (BC) using paired adapted and flat-ground motion commands. (c) Zero-shot whole-body teleoperation with onboard depth sensing and policy inference.
Method
VTR (%) ↑
Penetration (cm) ↓
Floating (cm) ↓
CP (%) ↑
Deviation (rad) ↓
Generation RTF ↓
Motions without hand contact
NEXUS (ours)
92.08
0.814
0.880
79.12
0.00984
1.646
TCRS (in PMT)
89.69
0.606
2.965
29.63
0.08884
8.038
Root-only
63.37
2.865
0.817
54.25
0.00000
0.260
Motions with hand contact
NEXUS (ours)
80.53
1.630
0.946
71.21
0.01473
1.864
TABLE I: Kinematic motion adaptation quality and efficiency.
Fig. 3: Baseline failures and NEXUS improvements at matching frames: (a) foot floating in both baselines; (b) foot penetration in Root-only and floating in TCRS; (c) hand penetration in both baselines; (d) TCRS distorts the non-contact pose by retracting the extended free leg. NEXUS improves contact while preserving the leg’s extension. Insets magnify circled regions; dashed circles mark the support foot.
Fig. 4: Whole-body control generalization compared with four open-source controllers. (a) NEXUS achieves the highest motion completion rates across the three dataset subsets, including ties. (b) NEXUS maintains the highest completion rates across terrain difficulty levels. (c) NEXUS attains the lowest body orientation error and highest foot-contact F1 on both flat ground and terrain. Tracking metrics use each controller’s own pre-failure window.
Metric
NEXUS
w/o vision
w/o T–S
SR (%) ↑
91.072±
14.959
87.771±
19.362
90.964±
15.938
Root Pos (cm) ↓
19.596±
05.398
20.089±
06.651
31.951±
04.472
MPKPE (cm) ↓
7.066±
02.949
7.603±
03.753
9.482±
03.533
MPKRE (rad) ↓
0.127±
00.010
0.128±
00.010
0.161±
00.015
Contact F1 (%) ↑
80.662±
02.545
80.523±
02.876
79.637±
02.859
TABLE II: Vision and teacher–student ablations.
Parameter
Value
Contact detection: entry / exit
Link-origin height [m]
0.18/0.25
Horizontal speed multiplier
0.20/0.30
Absolute vertical speed multiplier
0.15/0.25
Root horizontal speed floor [m/s]
1.0
Terrain queries and IK
TABLE III: Motion adaptation parameters.
Fig. 5: Motion adaptation schematics. (a) Height-based contact hysteresis, assuming both speed conditions remain satisfied. Shading denotes active contact. (b) PCHIP reconstruction of a scalar correction from contact keyframes, with constant extension beyond their span. Curves are illustrative, not measured results.
Reward term
Expression
Weight
Tracking
Root position
K0.3(∥Δp∥2)
0.5
Root orientation
K0.4(dR(R,Radapt)2)
0.5
Root linear velocity
K1.0(∥Δv∥2)
1.0
Root angular velocity
K3.0(∥Δω∥2)
1.0
Body position
K0.3(⟨∥Δpbrel∥2⟩b)
2.0
TABLE IV: Reward function details.
Randomized quantity
Distribution
Startup: teacher and student
Foot friction coefficient
U(0.3,1.2)
Torso CoM offset x [m]
U(−0.025,0.025)
Torso CoM offset y,z [m]
U(−0.05,0.05)
Joint encoder bias [rad]
U(−0.01,0.01)
Reset and pushes: teacher and student
TABLE V: Domain randomization applied during training.
Setting
Teacher (PPO)
Student (DAgger)
Training hyperparameters
Physics / control step [s]
0.005 / 0.020
0.005 / 0.020
Parallel environments
4×16,384
4×16,384
Initial learning rate
10−3
10−3
Learning-rate schedule
Adaptive (target KL 0.01 )
Constant
Rollout steps per environment
24
16
TABLE VI: Training hyperparameters and network architecture.
Humanoid behavior foundation models aim to acquire reusable whole-body control policies from broad human motion priors, enabling a single controller to produce diverse and expressive behaviors. However, existing motion-centric foundation policies largely assume that the reference motion is already physically compatible with the robot's surroundings. This assumption breaks when the demonstrator, operator, and robot inhabit different environments: a human motion may specify the intended behavior, but not the footholds, clearance, body height, or contact timing required by the robot's local terrain. We introduce \emph{Perceptive Behavior Foundation Model} (Perceptive BFM), a terrain-aware humanoid control framework that grounds human motion priors in robot-centric perception. The model preserves raw kinematic motion references as the behavioral interface, while using local terrain observations to adapt contacts, posture, and timing. To provide scalable terrain supervision, we develop \emph{terrain-conformal reference synthesis} (TCRS), which converts locomotion-oriented human motion clips into terrain-consistent references through contact-aware foothold construction, foot-geometry-aware swing optimization, support-aware root reconstruction, collision repair, and multi-point inverse kinematics. We then train a blind adapted-reference teacher and transfer its terrain-conformal behavior to a deployed raw-reference student through target-frame action alignment. The student is an identity-gated Transformer tracker whose terrain features enter through residual pathways initialized to preserve the motion-tracking prior and trained to produce local corrections only when needed.
Zifan Wang, Yizhao Li, Teli Ma +5
1Mondo Robotics · 2The Hong Kong University of Science and Technology (Guangzhou) · 4Artificial General Intelligence Institute, University of Science and Technology of China +2
Whole-body humanoid locomotion is challenging due to high-dimensional control, morphological instability, and the need for real-time adaptation to various terrains using onboard perception. Directly applying reinforcement learning (RL) with reward shaping to humanoid locomotion often leads to lower-body-dominated behaviors, whereas imitation-based RL can learn more coordinated whole-body skills but is typically limited to replaying reference motions without a mechanism to adapt them online from perception for terrain-aware locomotion. To address this gap, we propose a whole-body humanoid locomotion framework that combines skills learned from reference motions with terrain-aware adaptation. We first train a diffusion model on retargeted human motions for real-time prediction of terrain-aware reference motions. Concurrently, we train a whole-body reference tracker with RL using this motion data. To improve robustness under imperfectly generated references, we further fine-tune the tracker with a frozen motion generator in a closed-loop setting. The resulting system supports directional goal-reaching control with terrain-aware whole-body adaptation, and can be deployed on a Unitree G1 humanoid robot with onboard perception and computation. The hardware experiments demonstrate successful traversal over boxes, hurdles, stairs, and mixed terrain combinations. Quantitative results further show the benefits of incorporating online motion generation and fine-tuning the motion tracker for improved generalization and robustness.
Zewei Zhang, Kehan Wen, Michael Xu +7
Robotic Systems Lab, ETH Zurich · EPFL · Simon Fraser University +2
General-purpose motion trackers enable humanoid robots to follow diverse whole-body motions while maintaining balance, but are trained only on flat ground, failing to exploit bipedal mobility over complex terrain. Cross-terrain controllers, meanwhile, are task-specific or accept only low-dimensional locomotion commands. We introduce InterTrack, the first behavior world model (BWM) for robust whole-body tracking with environment interaction. Its Transformer jointly predicts the next action, state, and behavior distribution, learning environment-conditioned dynamics. To scale interaction training data, an automatic annotation pipeline reconstructs 3D support geometry from retargeted motions. At deployment, the policy handles commands implausible in the current environment in a "best-effort" manner. Quantitatively, InterTrack achieves an 81.3% success rate on terrain interaction (4.3 times the best evaluated baseline) and a 99.3% fall-recovery rate, while also improving free-space tracking and outperforming three leading tracking baselines across all of these regimes. To our knowledge, we provide the first demonstration of real-time cross-terrain whole-body teleoperation on a humanoid robot, alongside object interaction, stable responses to missing supports, and robust recovery from falls.
Ziyang Cheng, Tianshu Tang, Jinxin Lan +19
Tsinghua University · GigaAI · University of Shanghai for Science and Technology +3