Physics-Guided Residual Reinforcement Learning for Humanoid Narrow-Path Traversal
Authors: Tianchen Huang, Sisheng Chen, Wei Zhou, Haopeng Zhang, Jiarong Sun, Ya Wang, Yumin Wang, Deguang Lyu, +3 more
Organizations: Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
Traversing narrow paths is challenging for humanoid robots due to the sparse and safety-critical footholds required. Purely template-based or end-to-end reinforcement learning-based methods suffer from such harsh terrains. This paper proposes a two stage training framework for such narrow path traversing tasks, coupling a template-based foothold planner with a low-level foothold tracker from Stage-I training and a lightweight perception aided foothold modifier from Stage-II training. With the curriculum setup from flat ground to narrow paths across stages, the resulted controller in turn learns to robustly track and safely modify foothold targets to ensure precise foot placement over narrow paths. This framework preserves the interpretability from the physics-based template and takes advantage of the generalization capability from reinforcement learning, resulting in easy sim-to-real transfer. The learned policies outperform purely template-based or reinforcement learning-based baselines in terms of success rate, centerline adherence and safety margins. Validation on a Unitree G1 humanoid robot yields successful traversal of a 0.2m wide and 3m long beam for 20 trials without any failure.
Figures & tables
Fig. 1: Overview. The proposed framework uses 3D-LIPM as foothold planner, low-level policy from Stage-I training as foothold tracker and high-level policy from Stage-II training as foothold modifier for robust and safe narrow path walking. Successful experimental validation on the Unitree G1 humanoid robot is performed.
Fig. 2: The proposed framework for humanoid robot traversing narrow paths. A two-stage training curriculum is designed for the low-level foothold tracker and high-level foothold modifier. Different background colors indicate different operational frequencies
Fig. 3: Demonstration of foothold modification in the Stage-II training. The dashed polygon shows the initial foothold target from the foothold planner for the left leg, uinit(L) , while the solid polygons denote the final foothold targets from the foothold modifier for both legs, ufinal(L) (red) and ufinal(R) (blue). A smooth Gaussian penalty is given to the foothold’s lateral offset from the beam centerline, ∣yw−yw,center∣ , to encourage staying near the centerline
Method
Success rate (%)
Centerline dev. (m)
FP-RMSE (m)
No-Modifier
15
0.04690 ± 0.00057
0.01962 ± 0.00079
RL-Only
10
0.18192 ± 0.07075
— a
Ours
100
0.01639 ± 0.00117
0.02633 ± 0.00083
TABLE I: Baseline comparison results from walking on beams in simulation (mean ± std over 20 runs)
Configuration
Success rate (%)
Centerline dev. (m)
FP-RMSE (m)
w/o Stage-I disturbances
50
0.09696 ± 0.07876
0.05467 ± 0.05467
Ours
100
0.01639 ± 0.00117
0.02633 ± 0.00083
TABLE II: Ablation study results from walking on a 0.20m -wide beam in simulation (mean ± std over 20 runs)
Fig. 4: Top view of the anterior terrain height map based on the robot’s LiDAR sensor. The map spans x∈[0.1,1.1]m and y∈[−0.8,0.8]m with a resolution of 0.1m uniformly, providing a real-time terrain representation for foothold modification
Setting
Success rate (%)
Traversal rate (%)
BeamDojo [ 3 ] (G1, real beam)
80
88.16
Ours (G1, real beam)
100
100
TABLE III: Hardware comparison results from walking on a 3m×0.2m beam in reality (mean ± std). Ours: N=20 trials; BeamDojo: N=5 trials as reported
Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate robustly, safely, and accurately. While reinforcement learning with velocity-commanded policies has achieved remarkable robustness in humanoid locomotion, this approach lacks explicit control of the foothold placement, leading to unsafe behavior, such as stepping onto human feet, or imprecise navigation, hindering the following manipulation task. Conversely, explicit foothold-tracking policies offer a promising alternative by directly being commanded with target foot poses. However, existing approaches are often limited by unrealistic state assumptions, compromising real-world deployment, or they are part of staged pipelines, making them tied to specific downstream tasks. In this work, we introduce a novel, lightweight framework for training general-purpose 3D foothold-tracking policies. By dynamically providing footstep support through a goal sampler, this method enables the learned policy to be agnostic to specific terrains. Our new target representation effectively mitigates challenges arising in the real world, such as noisy and inaccurate pose estimation and foot contact estimation. Designed for direct real-world transfer, our policy acts as a standalone low-level controller that can be seamlessly paired with various high-level foothold generators. We demonstrate the effectiveness of our framework through extensive experiments in simulation and in the real world. By coupling our policy with different upstream planners, we achieve natural and accurate locomotion in challenging settings, paving the way for loco-manipulation tasks in complex environments.
Alessandro Montenegro, Shihao Li, Puze Liu +2
Politecnico di Milano · Tongji University · Technische Universität Darmstadt +2
Extending humanoid traversal to the open world is key to practical deployment in human environments, but remains challenging. The robot must use vision to ensure safe and reliable foot placement on heterogeneous terrain under highly dynamic motion, while producing coordinated, natural whole-body behaviors. We propose SSR, an efficient end-to-end framework for egocentric vision-based humanoid traversal that jointly learns these capabilities. SSR introduces imagined foothold guidance, which learns to model forthcoming swing-foot contacts and evaluates their support to guide pre-touchdown swings toward stable regions, reducing edge slips. It further employs equivariant latent-space symmetry augmentation to efficiently induce bilateral coordination under high-dimensional visual observations, and uses terrain-specific multi-discriminator motion priors to encourage human-like behavior across scenes. Extensive experiments show that SSR achieves safe, stable, and high-quality locomotion on diverse real-world terrains, including stairs with varied structures and extreme challenges such as wide gaps and high platforms, while enabling reliable long-horizon traversal in open outdoor environments.
We present a method for training reference-guided, perceptive reinforcement learning locomotion policies for humanoid robots in which reference trajectories are modulated in training to be consistent with terrain geometry. Aiming to deploy our method with standard navigation autonomy infrastructure, we synthesize SE(2)-controllable reference trajectories inside the RL training loop, projecting desired footsteps onto valid footholds and adjusting swing-foot and center-of-mass trajectories to match the terrain. The resulting policy exposes a clean SE(2) velocity interface compatible with standard navigation planners. In simulation, environmentally-conditioned references significantly improve reference tracking performance compared to environment agnostic references. On hardware, we integrate the policy with an MPC + control barrier function planner and demonstrate long-horizon (>70m) closed-loop autonomous navigation on the Unitree G1 through outdoor environments containing rough terrain and consecutive flights of stairs, with all sensing and computation onboard.
William D. Compton, Zachary Olkin, Aaron D. Ames
Department of Computing and Mathematical Sciences, California Institute of Technology, Pasadena, CA 91125