cs.ROOct 5, 2026

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

Authors: Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S. K., Samiran Datta, Abhay Dwivedi, Amit Shukla

Organizations: CAIR, Indian Institute of Technology Mandi, Mandi, India · SMME, Indian Institute of Technology Mandi, Mandi, India

Abstract

A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command (vx,vy,ωyaw)(v_x, v_y, ω_{yaw}) every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.

Figures & tables

Explore similar work

CardsList
  1. MARCH: Model-Assisted Reinforcement Learning for the Perceptive Control of Humanoids over Sparse Footholds

    Jun 9, 2026Codrin Crismariu, Ryan K. CosnerAgile LocomotionScalable Robot Learning

  2. asRoBallet: Closing the Sim2Real Gap via Friction-Aware Reinforcement Learning for Underactuated Spherical Dynamics

    Apr 27, 2026Fang Wan, Guangyi Huang, Tianyu Wu +5Scalable Robot LearningUnderactuated Dynamics

  3. Bicycle Acrobatics with Reinforcement Learning

    Aug 1, 2026Shamel Fahmi, Arianna Ilvonen, Samuel Zapolsky +8Scalable Robot LearningAgile Locomotion