Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Beijing University of Posts and Telecommunications · Beijing Jiaotong University
High-precision locomotion combines motion-command tracking with precise regulation of task-relevant physical states, enabling robots to interact reliably with their surroundings during motion. Joint end-to-end optimization can leave precision objectives insufficiently optimized, while reactive residual control adjusts actions only after deviations become observable. We present \textbf{LocoWM}, a world-model-guided preactive residual adaptation framework for high-precision locomotion. A base policy provides command-following locomotion, while an action-conditioned world model predicts a sequence of future physical states from proprioceptive history and the proposed base action. A residual adapter conditions on this predicted sequence to generate additive action corrections that compensate for anticipated deviations. Two-stage training first learns locomotion and action-conditioned dynamics, then freezes both modules while training the adapter, separating locomotion acquisition from precision adaptation. Experiments spanning terrain leveling, acceleration compensation, and push recovery demonstrate improved control precision and disturbance robustness over end-to-end and reactive residual baselines. Demos and code are available at: https://zhaozijie2022.github.io/LocoWM
Figures & tables
Figure 1: Left: Real-world transport of unsecured payloads across varied terrains. Center: LocoWM combines the base controller, action-conditioned world model, and residual adapter in a sequential pipeline: the base controller proposes atb , the world model predicts future task substates conditioned on atb , and the residual adapter uses these predictions to generate a preactive correction atr , yielding the executed action at=atb+atr . Right: Real-world demonstrations of (a) terrain leveling, (b) acceleration compensation, and (c) push recovery.
Figure 2: Overview of LocoWM. Stage 1: Base Policy & World Model. The base policy learns command-following locomotion, while the world model predicts future task substates from the same trajectories using a separate loss. Stage 2: Residual Adapter. The base policy and world model are frozen, and the adapter learns precision corrections. Inference. All modules are frozen; the base action, predicted substates, and residual correction are computed sequentially, yielding at=atb+atr .
Simulation Results
Pitch (deg)
Roll (deg)
∣az∣ (m/s 2 )
etrack
Succ (%)
Terrain
Method
RMS
peak
RMS
peak
mean
peak
Slope
Base Policy
28.18
29.98
9.96
21.51
1.58
8.54
0.28
45.3
End-to-End
15.18
32.38
2.99
6.15
0.89
7.19
0.42
62.8
React
5.39
20.89
2.15
8.16
0.83
7.12
0.43
59.8
Recon
3.01
11.97
1.25
7.14
0.74
5.94
0.46
61.5
LocoWM(ours)
2.61
3.90
0.31
1.23
0.41
3.10
0.45
91.1
Table 1: Terrain leveling across four terrains. Within each terrain, bold and underline mark the best and second best.
Figure 3: Acceleration compensation on flat ground. (a) Torso x -velocity tracking; (b) Acceleration–pitch responses; (c) Single-event tilt response.
Figure 4: Payload retention during push recovery. (a) Direction-averaged success rates versus push magnitude. (b) The 50% success-rate envelopes of all methods. (c) LocoWM success-rate map: angle denotes push direction, radius denotes force magnitude, and color denotes success rate.
Figure 5: Real-world terrain leveling on (a) a one-sided bridge and (b) bumps. In each panel, the top row shows successive video frames, while the bottom row combines robot cutouts extracted from individual frames onto a common background to illustrate the motion sequence.
Figure 6: Real-world acceleration compensation on flat ground.
Figure 7: Real-world push recovery with an unsecured bottle. Frames progress from left to right through force application, the transient response, and recovery; the bottle remains on the platform in both sequences.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Back-mounted platform stabilization. Left: Acceleration Compensation. Acceleration can destabilize a payload on a level platform; coordinated body tilting compensates for inertial loading. Right: Terrain Leveling. Uneven support heights disturb the platform orientation; adjustments to the leg configuration counteract torso tilt and maintain a level platform.
Term
Symbol
Group
Dim
Hist.
Noise
Scale
Velocity command
vc
P/C
3
6
–
1.0
Base angular velocity
ω
P/C/W
3
6
±0.2
0.25
Projected gravity
g
P/C/W
3
6
±0.05
1.0
Joint positions (legs)
q−q0
P/C
12
6
±0.01
1.0
Joint velocities
q˙
P/C
16
6
±1.5
0.05
Previous action
at−1
P/C
16
1
–
1.0
Appendix
Table 2: Observation terms. Group: P = policy, C = critic, W = world-model target.
Component
Expression
Weight (S1 / S2)
Lin. vel. tracking ( x )
exp(−(vxcmd−vx)2/σ2)
0.75 / 1.0
Lin. vel. tracking ( y )
exp(−(vycmd−vy)2/σ2)
0.75
Ang. vel. tracking ( z )
exp(−(ωzcmd−ωz)2/σ2)
0.75
Base angle tracking
exp(−∥gxy−gxy⋆∥2/σ2)
– / 0.5
Vertical velocity
vz2
-2.0 / -10.0
Roll/pitch rate
ωx2+ωy2
-0.05
Appendix
Table 3: Reward terms across the two training stages. A single value is shared by both stages.
Term
Min
Max
Unit
Op.
When
Base mass
−1.0
2.0
kg
add
startup
Body inertia
0.5
1.5
×
scale
startup
Base CoM offset ( x,y,z )
−0.05
0.05
m
add
startup
Foot static friction
0.5
1.0
–
abs
startup
Foot dynamic friction
0.5
0.8
–
abs
startup
Foot restitution
0.0
0.5
–
abs
startup
Appendix
Table 4: Domain randomization and perturbations. ∗ log-uniform.
Hyperparameter
Value
Optimizer
Adam
Initial learning rate
10−3
Learning-rate schedule
Adaptive, target KL 0.01
Rollout length per environment
24
Epochs per update
5
Discount factor
0.99
Appendix
Table 5: PPO hyperparameters used in both training stages.
Figure 9: Forward velocity tracking and pitch and roll responses during simulated terrain leveling.
Figure 10: Simulated terrain leveling with an upright payload.
Figure 11: Simulated acceleration compensation under step velocity commands.
Figure 12: Simulated acceleration compensation under triangular-wave velocity commands.
Figure 13: Simulated posture sequences during acceleration compensation under step commands in (a,b) and triangular-wave commands in (c,d). Insets show velocity tracking and normalized acceleration and pitch traces.
Figure 14: Simulated push recovery: forward velocity tracking (left) and payload pitch and roll (right).
Figure 15: Payload-retention success-rate maps for (a) End-to-End (E2E), (b) React, (c) Recon, and (d) LocoWM. Polar angle denotes push direction, radius denotes force magnitude, and color denotes success rate. Dashed black contours mark a 50% success rate.
Figure 16: Simulated push recovery under 150 N forces applied in eight body-frame directions. Frames progress from left to right through force application and recovery. Arrows indicate push direction, and inset traces show payload roll and pitch with the toppling thresholds.
Figure 17: Real-world terrain leveling with stacked blocks on (a) a slope and (b) a wave-shaped obstacle. Each panel shows successive video frames above and robot cutouts overlaid on a common background below.
Figure 18: Real-world push recovery under (a) backward, (b) forward, (c) leftward, and (d) rightward pushes. Each row progresses from force application to recovery, with the unsecured bottle retained on the platform.
Figure 19: Real-world acceleration compensation during (a) acceleration and (b) deceleration on flat ground. Successive frames from left to right show changes in leg configuration and platform pitch.
Task
Method
TrackLinErr (m/s) ↓
Grav-XY ↓
Success Rate (%) ↑
mean
std
mean
std
Command Track
Base-Policy
0.128
0.096
0.498
0.431
52.7
ReST-RL
0.123
0.108
0.083
0.212
91.0
LocoWM
0.120
0.102
0.063
0.165
94.1
Push Robot
Base-Policy
0.161
0.162
0.751
0.388
12.5
ReST-RL
0.150
0.148
0.128
0.281
80.9
Appendix
Table 6: Simulation results for humanoid tray transport. Bold entries mark the best mean error or success rate within each task.
Figure 20: Simulated humanoid transport during starting, walking, and stopping.
Figure 21: Simulated humanoid tray transport during (a) starting, (b) forward walking, and (c) stopping. Frames progress from left to right, with insets showing velocity tracking and payload orientation.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · The Chinese University of Hong Kong · Simplexity Robotics