Passive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion
Authors: Hyeonjin Choi, Joongheon Kim, Daekyum Kim
Organizations: School of Mechanical Engineering, Korea University, Seoul 02841, Republic of Korea · School of Electrical Engineering, Korea University, Seoul 02841, Republic of Korea · School of Smart Mobility, Korea University, Seoul 02841, Republic of Korea
Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagittal progression on flat collision geometry, complemented by curriculum-coupled reward terms. The core framework requires no reference trajectories, gait phases, or contact schedules. In a five-seed forward-locomotion study on a 29-DoF Unitree G1, the framework reduces mechanical cost of transport by 6.8-15.2% over commanded speeds of 0.5-2.0m/s without degrading velocity tracking. Mechanical-work decomposition attributes the reduction to positive actuator work, and reward-matched comparisons separate the guided regime's faster gait acquisition from the tilt's additional benefit to converged economy. The framework extends to unassisted omnidirectional locomotion, where its benefit persists once a walking-specific motion prior supplies kinematic coordination, the combination reducing speed-matched cost of transport by 18.7%. On hardware, forward cost of transport falls by 16.3% with the motion prior and by 4.5% without it, the latter within the trial-to-trial spread.
Figures & tables
Fig. 1: Overview of the proposed PDW-inspired slope-to-flat curriculum. Phase 1 : early policy search is conducted under a PDW-inspired tilted-gravity field, which supplies gravitational assistance to fore–aft step-to-step progression. Phase 2 : the tilt decays linearly, returning the training dynamics toward nominal. Phase 3 : all PDW-specific guidance is fully withdrawn, and the policy must sustain the induced gait structure under powered flat-ground command tracking. The collision geometry remains flat throughout; the depicted incline illustrates only the energetic effect of the tilted-gravity field.
Fig. 2: Training schedule of the proposed framework. The three curriculum components and the adversarial motion prior are shown against training time, divided into the three phases used throughout this article. Phase 1 : the tilted-gravity field supplies full slope-equivalent assistance, the PDW-inspired reward terms are active alongside the baseline objective, and commands are sagittally dominant. Phase 2 : the tilt decays linearly and the PDW-inspired reward terms are scaled down by the same curriculum factor σ(k) , while the command distribution expands toward the full omnidirectional set. Phase 3 : all PDW-specific guidance has vanished, and the policy is optimized under nominal flat-ground dynamics, the baseline objective alone, and the complete command distribution. The adversarial motion prior, when enabled, is active throughout training and is not part of the curriculum. Fig. 1 illustrates the mechanical intuition for the tilted-gravity component.
Fig. 3: Synchronized expansion of the translational command sectors and the yaw range over the curriculum. Left ( k<Kw ): commands are restricted to two sagittal sectors spanning ±45∘ about the forward and backward directions, and no yaw command is issued. Middle ( Kw≤k<Kt ): each sector expands continuously toward ±90∘ while the yaw range opens. Right ( k≥Kt ): the union of the two sectors covers the full 360∘ translational command range at the final yaw range. Green arrows denote admissible translational command directions; blue denotes the commanded yaw rate.
Condition
Tilted gravity
Velocity band
Positive-work
(VB)
penalty (PW)
Flat (baseline)
—
—
—
Flat + VB
—
✓
—
Flat + VB + PW
—
✓
✓
Slope + VB
✓
✓
—
Slope + VB + PW
✓
✓
✓
TABLE I: Conditions of the Forward-Locomotion Mechanism Study
Fig. 4: Cost of transport on the sagittal task across commanded speeds (mean over 5 seeds; error bars show the S.D. of total CoT). Total bar height gives the total CoT, split into a positive component (dark) and a negative component (light). The full curriculum (Slope + VB + PW) attains the lowest CoT at every speed and lowers it primarily by shrinking the positive component, while the negative component stays essentially unchanged across all conditions and speeds.
Fig. 5: Emergence of sustained stepping during training (mean ± s.d. over five seeds). Left: mean foot-contact events per episode. Right: fraction of episodes with at least five step events. The horizontal axis is the PPO policy-update index. The flat and tilted-gravity variants with matched reward terms follow similar trajectories, whereas the conventional baseline reaches sustained stepping substantially later.
θinit
W+/d
(W++W−)/d
vs. 5∘
Surv.
Failed
[J/m]
[J/m]
seeds
0∘
108.6±4.2
152.8±3.1
+4.9 %
0.982
0/3
3∘
108.1±0.5
151.6±1.1
+4.1 %
0.980
0/3
4∘
104.6±2.3
149.1±1.4
+2.4 %
0.975
0/3
5∘
102.2±0.5
145.6±1.9
—
0.992
0/3
6∘
104.8±1.7
148.9±3.0
+2.3 %
0.985
0/3
TABLE III: Initial-Tilt-Angle Sweep (3 Seeds)
Speed
CoT
CoT
Δ
Vel. err.
Gait sym.
(no tilt)
(perm. tilt)
Δ
Δ
0.5
0.4512
0.4882
+8.2%
+183.4 %
−32.8 %
1.0
0.4419
0.4554
+3.1%
+80.7 %
−4.8 %
1.5
0.4779
0.5388
+12.8%
+17.3 %
−3.8 %
2.0
0.5422
1.0848
+100.1%
+161.0 %
−8.0 %
TABLE IV: Permanent 5∘ Tilt vs. No Tilt, Rewards Identical (Paired Over 3 Seeds: 42, 123, 7)
Condition
CoT
Tracking
Stuck
(speed-matched)
error [m/s]
rate [%]
Baseline
0.5641±0.0072
0.110±0.005
0.02±0.03
Ours
0.5431±0.0196
0.118±0.009
0.07±0.12
Baseline + AMP
0.4982±0.0280
0.256±0.035
15.95±10.91
Ours + AMP
0.4588±0.0041
0.244±0.056
4.60±4.13
TABLE V: Forward-Command Evaluation under the Omnidirectional Task (Mean ± S.D., 3 Seeds)
Effect
Raw [%]
Speed-matched [%]
Full PDW-inspired curriculum, no AMP
−4.70±2.40
−3.75±2.68
Motion prior alone
−7.99±9.14
−11.67±5.30
Both
−17.03±2.62
−18.66±1.55
Additive prediction
−12.69±6.75
−15.42±2.64
Interaction
−4.34±4.13
−3.24±1.96
TABLE VI: 2×2 Decomposition of the CoT Effect (vs. Baseline; Mean ± S.D., 3 Seeds)
Fig. 6: Sagittal joint power over the gait cycle, from ipsilateral foot touchdown to the next touchdown of the same foot, for forward locomotion speed-matched over a realized-speed range of 0.50–1.40 m/s. Curves are means across three independently trained seeds and shaded bands denote ± one standard deviation across seeds; markers on the abscissa indicate each condition’s mean toe-off. Power is normalized by the nominal robot mass m . Note that the vertical scale differs between panels: the ankle operates at roughly one quarter of the hip and knee power range.
Fig. 7: Real-robot deployment on the Unitree G1. (a) Time-ordered snapshot sequences over roughly one walking stride for the four policies: Baseline and Ours without the motion prior, and Baseline + AMP and Ours + AMP. An overhead safety sling is attached in all trials and does not support the robot’s weight during walking. (b) Improvement of Ours over Baseline in cost of transport (solid) and in its positive-work component (hatched) by commanded direction, without AMP (green) and with AMP (red); positive values indicate greater efficiency for Ours, and distances are recovered from onboard sensing (Section IV-E ).
AMP
Direction
Baseline
Ours
ΔCoT
ΔCoT+
No
Forward
0.581
0.555
+4.5 %
+4.1 %
Backward
0.614
0.622
−1.3 %
−2.1 %
Lateral
0.587
0.585
+0.4 %
−1.9 %
Yes
Forward
0.521
0.436
+16.3 %
+22.3 %
Backward
0.499
0.499
+0.1 %
+3.0 %
Lateral
0.489
0.406
+17.0 %
+13.1 %
TABLE VIII: Hardware Cost of Transport by Direction (Measured Distance)
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Value
Simulation and control
Simulator
Isaac Lab 2.3.0 / Isaac Sim 5.1.0 (PhysX)
Physics step / decimation
0.005 s / 4
Control period Δt
0.02 s (50 Hz)
Episode length
20 s (1000 control steps)
Parallel environments
1024
Appendix
TABLE IX: Simulation, Actuation, and Domain Randomization
Item
Value
Recurrent core
single-layer LSTM, 256 units
MLP head hidden sizes
[256,128]
Observation stack
last 5 control steps
Actor input
96×5=480
Critic input
99×5=495 (adds base lin. vel.)
Learning rate
10−3 (adaptive), target KL 0.01
Appendix
TABLE X: Policy / Critic Network and PPO Hyperparameters
#
Term
Formula fi
Weight wi
1
Linear-velocity tracking
exp(−∥cv−vxyb∥2/σ2)
+1.0
2
Angular-velocity tracking (yaw)
exp(−(ωz∗−ωb,z)2/σ2)
+1.0
3
Alive bonus
1[not terminated]
+0.15
4
Base angular velocity (roll/pitch)
∥ωb,xy∥2
−0.05
5
Base orientation
∥gb,xy∥2
−1.0
6
Joint velocity
∑jq˙j2
−1×10−3
Appendix
TABLE XI: Baseline Locomotion Reward rbase , Identical Across All Conditions
Term / parameter
Value
Gate / meaning
Reward terms, each multiplied by σ(k)
Positive-work penalty
5×10−4
slope-gated
Velocity-band reward
0.3
slope-gated
lower threshold vmin
—
0.05 m/s
upper bound (forward study)
—
1.0 m/s, absolute
upper bound (omnidir.)
ηVB
1.2st∗
Appendix
TABLE XII: PDW-Inspired Reward Terms and Curriculum Schedule
Item
Value
Projected gravity / root height
3 / 1
Root lin. / ang. velocity (yaw frame)
3 / 3
Joint positions / velocities
29 / 29
Pelvis-relative key-body positions
12
Feature total
80
Discriminator MLP
160 – 1024 – 512 – 1
Appendix
TABLE XIII: AMP Motion Feature Φ (80-Dim) and Discriminator
TABLE XV: CoT by Commanded Direction (Mean ± S.D., 3 Seeds; n = per-seed episodes per bin)
Bin
Ours vs. Baseline
Ours + AMP vs. Baseline + AMP
forward
5.2
8.0
backward
2.5
20.6
lateral
2.9
17.1
walk_turn
3.3
14.7
all moving
3.6
14.0
Appendix
TABLE XVI: Directional CoT Reduction (%)
Fig. 8: Distribution of positive actuator work among the sagittal leg joints during forward locomotion, speed-matched over a realized-speed range of 0.50–1.40 m/s. Each robot bar sums to 100% within the three sagittal joint groups (Section IV-D7 ). Bars are means over three seeds, and the across-seed standard deviation is at most 2.9 percentage points for every segment. The hatched row labels the walking ranges reported by Farris and Sawicki [ 18 ] , with segment widths set by the normalized midpoints of those ranges; those three ranges are reported independently and need not sum to 100% , so the row is a set of reference intervals rather than a partition.