Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy are jointly trained with supervised force and velocity outputs and learned latent context. The estimated force conditions locomotion and additionally generates planar-velocity and yaw-rate commands for leash guidance through an analytical map. In sustained-force simulation sweeps, temporal means of componentwise force root mean square error range from 1.44 to 2.83,N. Compared with a domain-randomized baseline, the framework reduces velocity-tracking and base-orientation error scores by 21.6% and 46.5%, respectively, and increases mean survival from 68.29% to 94.60% in separate sustained-force tests. Unitree Go1 experiments demonstrate stationary vertical and horizontal force estimation, locomotion with an 8.5,kg payload whose weight exceeds the 70,N training force limit, and leash guidance using the same force-estimation interface.
Figures & tables
Fig. 1: Go1 demonstrations of the external-force interface: (a) payload-force estimation; (b) payload locomotion using online force estimates; (c,d) horizontal-force estimation; and (e,f) leash-guided locomotion. Arrows indicate applied forces.
Fig. 2: Shared force-estimation interface. Encoder Eψ and policy πθ are jointly trained with privileged critic Vϕ ; normalization is omitted here. Force estimates condition locomotion in both tasks and additionally generate leash commands through M . Payload commands are supplied externally. Deployed weights remain fixed.
Method
Online representation
Locomotion role
Guidance / compliance
Locomotion adaptation
RMA [ 7 ]
Latent context
Policy conditioning
—
BAS [ 13 ]
Mass, CoM shift, friction
Policy + safety
—
Beyond Robustness [ 14 ]
Load state and properties
Load stabilization
—
Force-based interaction
DeFazio et al. [ 10 ]
Tug estimate ( Δv proxy)
Policy conditioning
Discrete turn selection
TABLE I: Main deployment representations and their uses in locomotion and force interaction. Auxiliary velocity and latent outputs are omitted for force estimators.
Task rewards
Expression ( rj )
Weight ( wj )
Tracking lin. vel.
exp{−4vxycmd−vxy22}
1.0
Tracking ang. vel.
exp{−4(ωzcmd−ωz)2}
0.5
Regularization
Expression ( rj )
Weight ( wj )
Alive
1.0
0.5
Base height
(hbnominal−hb)2
−50.0
Orientation
gxy22
−5.0
TABLE II: Reward Function Components.
Quantity
Setting
Friction coefficient
U(0.5,1.25)
Velocity-push interval
15 s
Push velocity magnitude
U(0,1) m/s
Continuous-force resampling
Every 11 s
Continuous-force magnitude
U(0,70) N
Continuous-force direction
Uniform on S2
TABLE III: Domain randomization shared by all compared controllers.
Fig. 4: Stationary hardware force estimation. A 5 kg dumbbell is placed on the robot and removed. The curves show estimated force components, with payload gravity serving as the reference for the settled vertical value.
Fig. 5: Stationary horizontal-force estimation. Spring-scale magnitudes are 41 N at t=11 s for the x-axis pull and 45 N at t=24 s for the y-axis pull. The corresponding marked estimates are −39.7 N and −50.1 N; signs indicate the negative base-axis directions.
Fig. 6: Hardware locomotion with an 8.5 kg payload secured beneath the torso. Ours uses online force estimates from the proprioceptive estimator to condition the locomotion policy. Both controllers receive the same velocity commands.
Fig. 7: Recorded training rewards. Solid curves use Gaussian smoothing with standard deviation σsmooth=3 recorded samples; lighter traces are the unsmoothed recorded rewards. Oracle uses ground-truth conditioning.
Fig. 10: Hardware leash guidance using online force estimates. The sequence includes forward motion, a 180∘ turn, and small-radius turning. The heading reference is reconstructed offline from force estimates and measured yaw.
Method
X: 80 N
Y: 60 N
Z: 80 N
Oracle
100.00%
100.00%
84.08%
Ours
100.00%
100.00%
83.79%
IA
52.25%
84.77%
69.04%
DR
94.73%
44.04%
66.11%
TABLE IV: Simulation survival rates at the specified force conditions. Oracle is a privileged reference; bold marks the best among Ours, IA, and DR.
Humanoid robots operating in human-centered environments (e.g., homes, hospitals, and offices) must mitigate foot--ground impact transients, as impact-induced vibration and noise degrade user experience and repeated impacts accelerate hardware wear. However, existing low-noise locomotion training often relies on kinematic proxy objectives or fragile force sensors, and footwear-induced changes in contact dynamics introduce distribution shifts that hinder policy generalization.We present QuietWalk, a physics-informed reinforcement learning framework for ground-reaction-force-aware humanoid locomotion under diverse footwear conditions. QuietWalk employs an inverse-dynamics-constrained physics-informed neural network (PINN) to estimate per-foot vertical ground reaction forces (GRFs) from proprioceptive signals, and integrates the frozen predictor into the RL training loop to penalize predicted impact forces without requiring force sensors at deployment.On a held-out real-robot dataset, enforcing inverse-dynamics consistency reduces vertical GRF prediction errors by 82%-86% compared with a purely supervised predictor and improves the coefficient of determination from 0.39/0.67 to 0.99/0.99 for the left/right feet. On hardware at 1.2 m/s (barefoot; averaged over four floor materials), QuietWalk reduces mean A-weighted noise level by 7.17 dB and peak noise level by 4.98 dB under a consistent recording setup. Cross-footwear experiments (barefoot, skate shoes, athletic sneakers, and high heels) across multiple surfaces further demonstrate robust adaptation to footwear-induced contact variations.
Humanoid robots are entering our physical world at scale, yet as oversized toys--good at singing and dancing, but short on force-interaction capabilities for practical tasks. Bridging this gap necessitates prioritizing reliable contact perception as a fundamental requirement. Estimating external wrenches in humanoids is complicated by floating-base dynamics and indeterminate contact locations. Existing analytical frameworks require idealistic assumptions and hard-to-obtain measurements, which are often unavailable in practice. To bridge this gap, we propose SixthSense, a task-agnostic approach that infers whole-body contact timing, location, and wrenches from proprioception and IMU data alone. To capture the multi-modal dynamics between unstructured contact inputs and the uncertain motion outputs, we employ conditional flow matching to tokenize proprioceptive histories and estimate a spatiotemporally sparse contact-event flow. SixthSense serves as a plug-and-play perception module for applications including collision detection, physical human-robot interaction, and force-feedback teleoperation. Experiments across standing, walking, and whole-body motion-tracking policies showcased unprecedented performance in diverse behaviors.
Xingzhou Chen, Xiayan Xu, Yan Ning +8
The Hong Kong University of Science and Technology, Hong Kong SAR, China · Zhejiang University, Hangzhou, China · Tencent Robotics X, Shenzhen, China
In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry estimator for legged robots under unreliable contact conditions. The proposed estimator directly predicts relative displacement, relative rotation, and body-frame velocity from a recent history of onboard inertial and joint measurements. To improve robustness under unreliable contact conditions, we introduce a foot-aware cross-attention module that adaptively weights IMU and leg-wise kinematic tokens without relying on manually defined contact or slip thresholds. The estimator is trained with direct supervision and two physics-inspired auxiliary losses that promote kinematic consistency and reliable use of leg information. To reduce policy-specific overfitting and consequently improve sim-to-real transfer, simulation training incorporates policy randomization, followed by partial real-world fine-tuning of the temporal encoder and prediction head. Experiments across diverse indoor and outdoor terrains demonstrate consistent reductions in position drift compared with classical filtering-based, hybrid, and purely learning-based baselines. Ablation studies further validate the contributions of the proposed training objectives, policy randomization, and real-world fine-tuning, particularly under unreliable contacts and sim-to-real mismatch.
Taehyeon Kong, Woojin Kim, Jemin Hwangbo
Korea Advanced Institute of Science and Technology (KAIST), Yuseong-gu, Daejeon 34141, Republic of Korea