Passive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion
Authors: Hyeonjin Choi, Joongheon Kim, Daekyum Kim
Organizations: School of Mechanical Engineering, Korea University, Seoul 02841, Republic of Korea · School of Electrical Engineering, Korea University, Seoul 02841, Republic of Korea · School of Smart Mobility, Korea University, Seoul 02841, Republic of Korea
Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagittal progression on flat collision geometry, complemented by curriculum-coupled reward terms. The core framework requires no reference trajectories, gait phases, or contact schedules. In a five-seed forward-locomotion study on a 29-DoF Unitree G1, the framework reduces mechanical cost of transport by 6.8-15.2% over commanded speeds of 0.5-2.0m/s without degrading velocity tracking. Mechanical-work decomposition attributes the reduction to positive actuator work, and reward-matched comparisons separate the guided regime's faster gait acquisition from the tilt's additional benefit to converged economy. The framework extends to unassisted omnidirectional locomotion, where its benefit persists once a walking-specific motion prior supplies kinematic coordination, the combination reducing speed-matched cost of transport by 18.7%. On hardware, forward cost of transport falls by 16.3% with the motion prior and by 4.5% without it, the latter within the trial-to-trial spread.
Figures & tables
Fig. 1: Overview of the proposed PDW-inspired slope-to-flat curriculum. Phase 1 : early policy search is conducted under a PDW-inspired tilted-gravity field, which supplies gravitational assistance to fore–aft step-to-step progression. Phase 2 : the tilt decays linearly, returning the training dynamics toward nominal. Phase 3 : all PDW-specific guidance is fully withdrawn, and the policy must sustain the induced gait structure under powered flat-ground command tracking. The collision geometry remains flat throughout; the depicted incline illustrates only the energetic effect of the tilted-gravity field.
Fig. 2: Training schedule of the proposed framework. The three curriculum components and the adversarial motion prior are shown against training time, divided into the three phases used throughout this article. Phase 1 : the tilted-gravity field supplies full slope-equivalent assistance, the PDW-inspired reward terms are active alongside the baseline objective, and commands are sagittally dominant. Phase 2 : the tilt decays linearly and the PDW-inspired reward terms are scaled down by the same curriculum factor σ(k) , while the command distribution expands toward the full omnidirectional set. Phase 3 : all PDW-specific guidance has vanished, and the policy is optimized under nominal flat-ground dynamics, the baseline objective alone, and the complete command distribution. The adversarial motion prior, when enabled, is active throughout training and is not part of the curriculum. Fig. 1 illustrates the mechanical intuition for the tilted-gravity component.
Fig. 3: Synchronized expansion of the translational command sectors and the yaw range over the curriculum. Left ( k<Kw ): commands are restricted to two sagittal sectors spanning ±45∘ about the forward and backward directions, and no yaw command is issued. Middle ( Kw≤k<Kt ): each sector expands continuously toward ±90∘ while the yaw range opens. Right ( k≥Kt ): the union of the two sectors covers the full 360∘ translational command range at the final yaw range. Green arrows denote admissible translational command directions; blue denotes the commanded yaw rate.
Condition
Tilted gravity
Velocity band
Positive-work
(VB)
penalty (PW)
Flat (baseline)
—
—
—
Flat + VB
—
✓
—
Flat + VB + PW
—
✓
✓
Slope + VB
✓
✓
—
Slope + VB + PW
✓
✓
✓
TABLE I: Conditions of the Forward-Locomotion Mechanism Study
Fig. 4: Cost of transport on the sagittal task across commanded speeds (mean over 5 seeds; error bars show the S.D. of total CoT). Total bar height gives the total CoT, split into a positive component (dark) and a negative component (light). The full curriculum (Slope + VB + PW) attains the lowest CoT at every speed and lowers it primarily by shrinking the positive component, while the negative component stays essentially unchanged across all conditions and speeds.
Fig. 5: Emergence of sustained stepping during training (mean ± s.d. over five seeds). Left: mean foot-contact events per episode. Right: fraction of episodes with at least five step events. The horizontal axis is the PPO policy-update index. The flat and tilted-gravity variants with matched reward terms follow similar trajectories, whereas the conventional baseline reaches sustained stepping substantially later.
θinit
W+/d
(W++W−)/d
vs. 5∘
Surv.
Failed
[J/m]
[J/m]
seeds
0∘
108.6±4.2
152.8±3.1
+4.9 %
0.982
0/3
3∘
108.1±0.5
151.6±1.1
+4.1 %
0.980
0/3
4∘
104.6±2.3
149.1±1.4
+2.4 %
0.975
0/3
5∘
102.2±0.5
145.6±1.9
—
0.992
0/3
6∘
104.8±1.7
148.9±3.0
+2.3 %
0.985
0/3
TABLE III: Initial-Tilt-Angle Sweep (3 Seeds)
Speed
CoT
CoT
Δ
Vel. err.
Gait sym.
(no tilt)
(perm. tilt)
Δ
Δ
0.5
0.4512
0.4882
+8.2%
+183.4 %
−32.8 %
1.0
0.4419
0.4554
+3.1%
+80.7 %
−4.8 %
1.5
0.4779
0.5388
+12.8%
+17.3 %
−3.8 %
2.0
0.5422
1.0848
+100.1%
+161.0 %
−8.0 %
TABLE IV: Permanent 5∘ Tilt vs. No Tilt, Rewards Identical (Paired Over 3 Seeds: 42, 123, 7)
Condition
CoT
Tracking
Stuck
(speed-matched)
error [m/s]
rate [%]
Baseline
0.5641±0.0072
0.110±0.005
0.02±0.03
Ours
0.5431±0.0196
0.118±0.009
0.07±0.12
Baseline + AMP
0.4982±0.0280
0.256±0.035
15.95±10.91
Ours + AMP
0.4588±0.0041
0.244±0.056
4.60±4.13
TABLE V: Forward-Command Evaluation under the Omnidirectional Task (Mean ± S.D., 3 Seeds)
Effect
Raw [%]
Speed-matched [%]
Full PDW-inspired curriculum, no AMP
−4.70±2.40
−3.75±2.68
Motion prior alone
−7.99±9.14
−11.67±5.30
Both
−17.03±2.62
−18.66±1.55
Additive prediction
−12.69±6.75
−15.42±2.64
Interaction
−4.34±4.13
−3.24±1.96
TABLE VI: 2×2 Decomposition of the CoT Effect (vs. Baseline; Mean ± S.D., 3 Seeds)
Fig. 6: Sagittal joint power over the gait cycle, from ipsilateral foot touchdown to the next touchdown of the same foot, for forward locomotion speed-matched over a realized-speed range of 0.50–1.40 m/s. Curves are means across three independently trained seeds and shaded bands denote ± one standard deviation across seeds; markers on the abscissa indicate each condition’s mean toe-off. Power is normalized by the nominal robot mass m . Note that the vertical scale differs between panels: the ankle operates at roughly one quarter of the hip and knee power range.
Fig. 7: Real-robot deployment on the Unitree G1. (a) Time-ordered snapshot sequences over roughly one walking stride for the four policies: Baseline and Ours without the motion prior, and Baseline + AMP and Ours + AMP. An overhead safety sling is attached in all trials and does not support the robot’s weight during walking. (b) Improvement of Ours over Baseline in cost of transport (solid) and in its positive-work component (hatched) by commanded direction, without AMP (green) and with AMP (red); positive values indicate greater efficiency for Ours, and distances are recovered from onboard sensing (Section IV-E ).
AMP
Direction
Baseline
Ours
ΔCoT
ΔCoT+
No
Forward
0.581
0.555
+4.5 %
+4.1 %
Backward
0.614
0.622
−1.3 %
−2.1 %
Lateral
0.587
0.585
+0.4 %
−1.9 %
Yes
Forward
0.521
0.436
+16.3 %
+22.3 %
Backward
0.499
0.499
+0.1 %
+3.0 %
Lateral
0.489
0.406
+17.0 %
+13.1 %
TABLE VIII: Hardware Cost of Transport by Direction (Measured Distance)
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Value
Simulation and control
Simulator
Isaac Lab 2.3.0 / Isaac Sim 5.1.0 (PhysX)
Physics step / decimation
0.005 s / 4
Control period Δt
0.02 s (50 Hz)
Episode length
20 s (1000 control steps)
Parallel environments
1024
Appendix
TABLE IX: Simulation, Actuation, and Domain Randomization
Item
Value
Recurrent core
single-layer LSTM, 256 units
MLP head hidden sizes
[256,128]
Observation stack
last 5 control steps
Actor input
96×5=480
Critic input
99×5=495 (adds base lin. vel.)
Learning rate
10−3 (adaptive), target KL 0.01
Appendix
TABLE X: Policy / Critic Network and PPO Hyperparameters
#
Term
Formula fi
Weight wi
1
Linear-velocity tracking
exp(−∥cv−vxyb∥2/σ2)
+1.0
2
Angular-velocity tracking (yaw)
exp(−(ωz∗−ωb,z)2/σ2)
+1.0
3
Alive bonus
1[not terminated]
+0.15
4
Base angular velocity (roll/pitch)
∥ωb,xy∥2
−0.05
5
Base orientation
∥gb,xy∥2
−1.0
6
Joint velocity
∑jq˙j2
−1×10−3
Appendix
TABLE XI: Baseline Locomotion Reward rbase , Identical Across All Conditions
Term / parameter
Value
Gate / meaning
Reward terms, each multiplied by σ(k)
Positive-work penalty
5×10−4
slope-gated
Velocity-band reward
0.3
slope-gated
lower threshold vmin
—
0.05 m/s
upper bound (forward study)
—
1.0 m/s, absolute
upper bound (omnidir.)
ηVB
1.2st∗
Appendix
TABLE XII: PDW-Inspired Reward Terms and Curriculum Schedule
Item
Value
Projected gravity / root height
3 / 1
Root lin. / ang. velocity (yaw frame)
3 / 3
Joint positions / velocities
29 / 29
Pelvis-relative key-body positions
12
Feature total
80
Discriminator MLP
160 – 1024 – 512 – 1
Appendix
TABLE XIII: AMP Motion Feature Φ (80-Dim) and Discriminator
TABLE XV: CoT by Commanded Direction (Mean ± S.D., 3 Seeds; n = per-seed episodes per bin)
Bin
Ours vs. Baseline
Ours + AMP vs. Baseline + AMP
forward
5.2
8.0
backward
2.5
20.6
lateral
2.9
17.1
walk_turn
3.3
14.7
all moving
3.6
14.0
Appendix
TABLE XVI: Directional CoT Reduction (%)
Fig. 8: Distribution of positive actuator work among the sagittal leg joints during forward locomotion, speed-matched over a realized-speed range of 0.50–1.40 m/s. Each robot bar sums to 100% within the three sagittal joint groups (Section IV-D7 ). Bars are means over three seeds, and the across-seed standard deviation is at most 2.9 percentage points for every segment. The hatched row labels the walking ranges reported by Farris and Sawicki [ 18 ] , with segment widths set by the normalized midpoints of those ranges; those three ranges are reported independently and need not sum to 100% , so the row is a set of reference intervals rather than a partition.
Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective. We instead treat these functions through distinct mechanisms: rewards for task specification, constraints for operational limits, energy minimization for gait preference, and exteroceptive perception for adapting energy use to terrain difficulty. We show that these components jointly enable efficient, terrain-adaptive locomotion, and that removing each component exposes a distinct failure mode. Our formulation removes explicit gait priors (including air-time, contact-count, and foot-clearance targets) in favor of emergent behavior. Compared to a conventional complex-reward baseline, our formulation achieves comparable terrain traversal while reducing cost of transport by 56% and operational-limit violations by 96%. The resulting policies transfer zero-shot to a physical Unitree Go2 using LiDAR-based elevation mapping. Project website with videos: https://tinyurl.com/locomposition.
Loukas Kordos, Leonard T. Franz, Simon Rappenecker +4
Technical University of Munich. · University of Tübingen. · Hertie Institute for Clinical Brain Research & Center for Integrative Neuroscience. +1
We propose a unified reinforcement learning framework that enables a single policy to perform walking, running, and fall recovery on the Unitree G1 humanoid robot, validated on physical hardware without any explicit mode-switching command at deployment. The framework extends Adversarial Motion Priors (AMP) by replacing the conventional global reference distribution with a state-dependent gate that routes each training transition to one of two discriminators: a dedicated recovery discriminator and a velocity-conditioned locomotion discriminator that jointly covers walking and running. The gate is defined by a single fixed threshold on projected gravity: the recovery discriminator is activated when body tilt exceeds approximately 37∘ from vertical (∣gz+1∣>0.6); otherwise the locomotion discriminator is used, with the normalized commanded velocity serving as a condition that selects the appropriate reference trajectory between walk and run clips. Only three LAFAN1 reference clips are required to regularize the complete behavior set. At deployment, a single frozen ONNX policy executes at 50,Hz with no runtime mode logic; hardware experiments demonstrate successful recovery from both prone and supine falls and smooth walk-to-run transitions under the same controller.
Yidan Lu, Yichao Zhong, Liu Zhao +2
The University of Hong Kong, Hong Kong SAR, China.
Humanoid robots operating in human-centered environments (e.g., homes, hospitals, and offices) must mitigate foot--ground impact transients, as impact-induced vibration and noise degrade user experience and repeated impacts accelerate hardware wear. However, existing low-noise locomotion training often relies on kinematic proxy objectives or fragile force sensors, and footwear-induced changes in contact dynamics introduce distribution shifts that hinder policy generalization.We present QuietWalk, a physics-informed reinforcement learning framework for ground-reaction-force-aware humanoid locomotion under diverse footwear conditions. QuietWalk employs an inverse-dynamics-constrained physics-informed neural network (PINN) to estimate per-foot vertical ground reaction forces (GRFs) from proprioceptive signals, and integrates the frozen predictor into the RL training loop to penalize predicted impact forces without requiring force sensors at deployment.On a held-out real-robot dataset, enforcing inverse-dynamics consistency reduces vertical GRF prediction errors by 82%-86% compared with a purely supervised predictor and improves the coefficient of determination from 0.39/0.67 to 0.99/0.99 for the left/right feet. On hardware at 1.2 m/s (barefoot; averaged over four floor materials), QuietWalk reduces mean A-weighted noise level by 7.17 dB and peak noise level by 4.98 dB under a consistent recording setup. Cross-footwear experiments (barefoot, skate shoes, athletic sneakers, and high heels) across multiple surfaces further demonstrate robust adaptation to footwear-induced contact variations.