Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy are jointly trained with supervised force and velocity outputs and learned latent context. The estimated force conditions locomotion and additionally generates planar-velocity and yaw-rate commands for leash guidance through an analytical map. In sustained-force simulation sweeps, temporal means of componentwise force root mean square error range from 1.44 to 2.83,N. Compared with a domain-randomized baseline, the framework reduces velocity-tracking and base-orientation error scores by 21.6% and 46.5%, respectively, and increases mean survival from 68.29% to 94.60% in separate sustained-force tests. Unitree Go1 experiments demonstrate stationary vertical and horizontal force estimation, locomotion with an 8.5,kg payload whose weight exceeds the 70,N training force limit, and leash guidance using the same force-estimation interface.
Figures & tables
Fig. 1: Go1 demonstrations of the external-force interface: (a) payload-force estimation; (b) payload locomotion using online force estimates; (c,d) horizontal-force estimation; and (e,f) leash-guided locomotion. Arrows indicate applied forces.
Fig. 2: Shared force-estimation interface. Encoder Eψ and policy πθ are jointly trained with privileged critic Vϕ ; normalization is omitted here. Force estimates condition locomotion in both tasks and additionally generate leash commands through M . Payload commands are supplied externally. Deployed weights remain fixed.
Method
Online representation
Locomotion role
Guidance / compliance
Locomotion adaptation
RMA [ 7 ]
Latent context
Policy conditioning
—
BAS [ 13 ]
Mass, CoM shift, friction
Policy + safety
—
Beyond Robustness [ 14 ]
Load state and properties
Load stabilization
—
Force-based interaction
DeFazio et al. [ 10 ]
Tug estimate ( Δv proxy)
Policy conditioning
Discrete turn selection
TABLE I: Main deployment representations and their uses in locomotion and force interaction. Auxiliary velocity and latent outputs are omitted for force estimators.
Task rewards
Expression ( rj )
Weight ( wj )
Tracking lin. vel.
exp{−4vxycmd−vxy22}
1.0
Tracking ang. vel.
exp{−4(ωzcmd−ωz)2}
0.5
Regularization
Expression ( rj )
Weight ( wj )
Alive
1.0
0.5
Base height
(hbnominal−hb)2
−50.0
Orientation
gxy22
−5.0
TABLE II: Reward Function Components.
Quantity
Setting
Friction coefficient
U(0.5,1.25)
Velocity-push interval
15 s
Push velocity magnitude
U(0,1) m/s
Continuous-force resampling
Every 11 s
Continuous-force magnitude
U(0,70) N
Continuous-force direction
Uniform on S2
TABLE III: Domain randomization shared by all compared controllers.
Fig. 4: Stationary hardware force estimation. A 5 kg dumbbell is placed on the robot and removed. The curves show estimated force components, with payload gravity serving as the reference for the settled vertical value.
Fig. 5: Stationary horizontal-force estimation. Spring-scale magnitudes are 41 N at t=11 s for the x-axis pull and 45 N at t=24 s for the y-axis pull. The corresponding marked estimates are −39.7 N and −50.1 N; signs indicate the negative base-axis directions.
Fig. 6: Hardware locomotion with an 8.5 kg payload secured beneath the torso. Ours uses online force estimates from the proprioceptive estimator to condition the locomotion policy. Both controllers receive the same velocity commands.
Fig. 7: Recorded training rewards. Solid curves use Gaussian smoothing with standard deviation σsmooth=3 recorded samples; lighter traces are the unsmoothed recorded rewards. Oracle uses ground-truth conditioning.
Fig. 10: Hardware leash guidance using online force estimates. The sequence includes forward motion, a 180∘ turn, and small-radius turning. The heading reference is reconstructed offline from force estimates and measured yaw.
Method
X: 80 N
Y: 60 N
Z: 80 N
Oracle
100.00%
100.00%
84.08%
Ours
100.00%
100.00%
83.79%
IA
52.25%
84.77%
69.04%
DR
94.73%
44.04%
66.11%
TABLE IV: Simulation survival rates at the specified force conditions. Oracle is a privileged reference; bold marks the best among Ours, IA, and DR.