Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human motions. Yet these capabilities remain largely confined to flat ground, in part because motion datasets rarely include aligned terrain geometry and because retargeting terrain interactions to complex musculoskeletal bodies is challenging. We present TERRA, an end-to-end pipeline for terrain-aware retargeting and control of musculoskeletal locomotion. From kinematic trajectories alone, TERRA combines terrain priors, estimated contacts, and negative free-space evidence to recover task-relevant support geometry. TERRA further considers anatomical, tendon-continuity, and contact constraints during retargeting. Using the resulting motion-terrain pairs from five datasets, we successfully train a single muscle-actuated control policy on 9.4 hours of diverse locomotion. Across reconstruction, retargeting, and held-out tracking benchmarks, TERRA improves terrain accuracy, sharply reduces anatomical and interaction violations, and achieves the highest observed completion rate over supported terrain families. Overall, TERRA provides a practical route from scene-less motion data to muscle-actuated locomotion over diverse non-flat terrain. Project website: https://cnai.epfl.ch/terra/
Figures & tables
Method
Kinematic constraints
Terrain reconstruction
Terrain retargeting
MSK
GMR [ 24 ]
Joint bounds
✗
✗
✗
OmniRetarget [ 27 ]
Hard interaction
✗
✓
✗
SceneBot [ 21 ]
Joint bounds
✓
✓
✗
MuscleMimic [ 7 ]
Joint/equality
✗
✗
✓
TERRA
Hard + anatomy
✓
✓
✓
TABLE I: Capabilities of representative motion-processing methods. MSK denotes support for musculoskeletal target models.
Fig. 2: TERRA method overview. Heterogeneous motion data are converted to SMPL-H, used to reconstruct compact collision terrain, and retargeted to the musculoskeletal model with terrain-aware constraints. The resulting paired trajectory–terrain scenes train a muscle-actuated tracking policy, whose biomechanical outputs are compared with subject-matched force-plate and EMG data on held-out motions.
Fig. 3: Terrain-reconstruction. Method comparison on a Vielemeyer 10∘ descent ramp (top) and PRISM stepping boxes (bottom). Black crosses denote the same reference toe contacts in every panel, and dotted segments connect each contact to the reconstructed height at the same horizontal location. Vertical height is exaggerated for visibility.
(a) Known terrain
All known terrain
Gait120
Darmstadt
Vielemeyer
Method
Family acc. (%) ↑ ( N=5,474 )
Grade MAE ( ∘ ) ↓ ( N=1190 )
Riser MAE (mm) ↓ ( N=1184 )
Stool MAE (mm) ↓ ( N=1102 )
Riser MAE (mm) ↓ ( N=1282 )
Grade MAE ( ∘ ) ↓ ( N=568 )
Contact least squares
N/A
1.683 ± 3.533
23.42 ± 8.71
490.00 ± 0.00
150.49 ± 57.41
1.199 ± 1.088
Voronoi
N/A
N/A
42.24 ± 39.54
80.90 ± 121.19
24.81 ± 36.39
N/A
TERRA (without physical cues)
73.5
0.220 ± 0.235
12.25 ± 5.35
45.62 ± 13.57
7.12 ± 6.10
0.536 ± 0.333 (N=566)
TERRA
99.9
0.223 ± 0.244
8.08 ± 6.99
45.62 ± 13.57
5.17 ± 4.36
0.536 ± 0.353
TABLE II: Terrain classification, reconstruction, and motion–terrain consistency across datasets. Entries are per-motion mean ± population standard deviation.
Retargeting
Policy tracking
Method
Success
Dten (%) ↓
Dcol (%) ↓
Dpen (%) ↓
Dskat (%) ↓
Dfloat (%) ↓
Contact F1 (%) ↑
Success (%) ↑
MPJPE (mm) ↓
MM-MoCap-Body
6478/6498
2.208 ± 3.852
12.09 ± 19.90
54.56 ± 35.53
4.86 ± 5.10
22.80 ± 36.78
92.86 ± 7.47
47.34 ± 0.22
90.26 ± 0.60
GMR
6498/6498
0.556 ± 1.087
21.82 ± 39.14
19.98 ± 24.19
9.25 ± 14.62
11.82 ± 11.27
74.19 ± 21.01
56.43 ± 0.64
91.97 ± 1.73
OmniRetarget
6392/6498
0.466 ± 1.243
4.98 ± 10.82
31.97 ± 29.22
12.70 ± 10.54
3.61 ± 9.79
86.05 ± 11.79
80.29 ± 1.39
87.23 ± 2.92
TERRA (no residuals)
6498/6498
0.134 ± 0.407
3.93 ± 8.45
33.45 ± 33.63
8.88 ± 7.03
4.18 ± 11.00
88.49 ± 9.51
86.52 ± 0.25
80.58 ± 2.31
TERRA+Voronoi
6470/6498
0.081 ± 0.285
0.54 ± 1.35
2.99 ± 14.97
2.63 ± 2.82
9.97 ± 18.73
92.96 ± 5.96
82.16 ± 0.41
71.65 ± 1.00
TABLE III: Retargeting performance across datasets and held-out policy-tracking performance.
Parameter
Value
Stationary contacts
vs=0.30 m/s; nmin=5 ; We=±0.30 s; ϵloc=0.05 m.
Geometry margins
Contact expansion 0.10/0.05 m; maximum extension 0.00/0.15 m; free-space clearance 0.03/0.02 m (vertical/horizontal).
Deploying humanoid robots in unstructured terrain remains an open problem. While classic reinforcement learning struggles with the sheer complexity of real-world interactions, more promising methods leveraging human priors remain limited to models lacking contextual awareness. The restricted motion synthesis is a direct consequence of existing dataset pipelines failing to capture human-scene sequences in challenging environments. To bridge this gap between humanoid learning and scene reconstruction, we introduce the Egocentric Human-Terrain Reconstruction (EgoHTR) dataset. We develop and open-source a reconstruction pipeline capturing 55 scene-aligned 4D human motion sequences in diverse, complex environments using a multi-sensor setup of egocentric wearables and a portable 3D scanner. The resulting dataset comprises over 150k frames, which we evaluate against motion-capture ground truth, demonstrating state-of-the-art accuracy and establishing a rigorous benchmark for human motion analysis and synthesis. Further, we leverage this data to train perceptive locomotion policies, demonstrating hardware deployment on a Unitree G1 for reconstructed reference motions. Our pipeline enables community-driven dataset extensions and factors the problem to help researchers build foundational, context-aware robots that reliably traverse uneven terrain.
Alex Brandes, Haig Conti Georges Sajelian, Manthan Patel +11
Retargeting human kinematic reference motion onto a robot's morphology remains a formidable challenge. Existing methods often produce physical inconsistencies, such as foot sliding, self-collisions, or dynamically infeasible motions, which hinder downstream imitation learning. We propose a bilevel optimization framework that jointly adapts reference motions to a robot's morphology while training a tracking policy using reinforcement learning. To make the optimization tractable, we derive an approximate gradient for the upper-level loss. Our framework requires only a sparse set of semantic rigid-body correspondences and eliminates the need for manual tuning by identifying optimal values for a parameterization expressive enough to preserve characteristic motion across different embodiments. Moreover, by integrating retargeting directly with physics simulation, we produce physically plausible motions that facilitate robust imitation learning. We validate our method in simulation and on hardware, demonstrating challenging motions for morphologies that differ significantly from a human, including retargeting onto a quadruped.
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracking cannot resolve the physical ambiguities of interacting with objects and uneven terrain. To address this, we introduce SceneBot, a unified motion-tracking framework capable of handling freespace locomotion, terrain traversal, and whole-body manipulation. SceneBot conditions a single policy on both reference motions and per-link contact labels, explicitly defining expected environmental interactions. To overcome the lack of annotated interaction data, we propose a hindsight scene reconstruction approach that infers scene-interaction graphs from retargeted human motion. Trained on 7.5 hours of this reconstructed, contact-rich data, SceneBot successfully generalizes to unseen motions and environments. Our results demonstrate that SceneBot is the first general framework to seamlessly unify free-space and contact-rich behaviors executing complex, long-horizon tasks like carrying a box upstairs and establishing contact conditioning as a powerful interface for humanoid control. All code and data will be open-sourced. More demos and information are available at: https://ericcsr.github.io/scenebot/
Sirui Chen, Shibo Zhao, Zhen Wu +3
Stanford, Amazon FAR United States · Amazon FAR United States · CMU, Amazon FAR United States