Quadruped robots are increasingly expected to carry objects while moving through human environments. But what happens when a person interacts directly with the payload rather than with the robot? If the payload is unrestrained, the robot must distinguish intentional external interactions from ordinary payload motion, while still keeping the load balanced and maintaining stable locomotion. How can a quadruped infer and compliantly respond to such interactions using only onboard measurements? In this work, we develop a force-aware locomotion framework that treats payload interactions as commands that shape the motion of the combined robot-payload system. Our approach separates the learning of force-aware locomotion and force estimation on an unrestrained payload. We combine a compliant load-carrying policy with a causal force estimator, trained through estimator-in-the-loop data aggregation and finetuning, to predict interactions from onboard robot measurements. Our simulations and real-world experiments show that the resulting controller can maintain stable payload-carrying locomotion, yield compliantly to external interactions, and use the inferred force to support human-guided changes in the robot's trajectory.
Figures & tables
Figure 1: Robot inferring and reacting to forces felt on an unrestrained payload. (Top) The robot detects a payload slip and tries to balance the payload, before returning to the planned path. (Bottom) The robot detects an external interaction on the payload, compliantly yields to the interaction, and updates its path to avoid a potential hazard.
Figure 2: Overview of our proposed approach for compliant locomotion under payload interactions. (Left) Privileged policy learning trains a force-aware locomotion policy, while a recurrent adaptation module learns to recover its privileged latent from deployable observations. (Center) A causal payload force estimator predicts interaction activity, direction, and magnitude from onboard measurements and is trained through mixed-distribution data aggregation. (Right) The estimator and adaptation module are frozen while the actor is finetuned under estimator-driven rollouts using PPO and behavioral anchoring toward the frozen oracle force-adapted policy (privileged policy). Green components require privileged or true force inputs, whereas purple components form the deployable inference path.
Method
ρcmd↑
eθ (deg) ↓
Termination (%) ↓
Oracle
0.865±0.001
1.77±0.04
19.13±0.31
Facet
0.252±0.002
8.99±0.21
52.91±0.23
Facet-L
0.702±0.002
7.27±0.16
21.57±0.08
Ours-Analytic
0.602±0.003
3.37±0.15
33.48±0.44
Ours
0.826±0.003
3.73±0.15
19.84±0.24
Table I: Aggregate compliance response and full-grid termination. Values in the columns represent the mean and std dev across five seeds.
Figure 3: Compliance across diverse force and impedance settings. (left) Commanded compliance ratio at Kp=16 N/m and mL=6 kg. (right) Displacement under an 8 N force as commanded stiffness varies. Curves represent the means and shaded regions show standard deviations across five evaluation seeds.
Force input
Magnitude ratio ↑
Direction error (deg) ↓
Detection (%) ↑
Onset (s) ↓
Learned
0.886±0.003
10.87±0.34
98.90±0.29
0.175±0.001
Analytic
0.445±0.002
14.24±0.47
86.84±0.44
0.433±0.006
Table II: Quantitative-analysis of the force estimator. Values in the columns represent the mean and std dev across five seeds
Method
Clearance (%) ↑
Joint success (%) ↑
Goal error (m) ↓
Oracle
92.4±0.8
92.4±0.7
0.038±0.002
Facet-L
0.0±0.0
0.0±0.0
0.093±0.002
Ours-Analytic
34.8±1.8
34.3±1.7
0.055±0.003
Ours
77.1±1.6
76.9±1.5
0.068±0.002
Table III: Force-guided obstacle-avoidance performance. Values in the columns show the mean and std. dev. across five seeds.
Method
vx (m/s)
∣vy∣ (m/s)
vxy (m/s)
Mean path deviation (m) ↓
Goal error (m) ↓
Mean tilt (deg) ↓
Peak tilt (deg) ↓
Straight Line
Facet -L
0.259±0.001
0.001±0.001
0.259±0.001
0.025±0.017
0.026±0.025
2.05±0.43
6.14±1.91
Ours
0.252±0.002
0.003±0.003
0.252±0.002
0.045±0.020
0.042±0.015
1.70±0.82
7.40±5.01
Human Guided
Facet -L
0.232±0.006
0.109±0.023
0.258±0.009
0.058±0.017
0.089±0.025
2.07±0.43
8.38±2.13
Ours
0.195±0.012
0.160±0.032
0.254±0.018
0.096±0.015
0.045±0.045
3.45±0.53
13.37±4.62
Table IV: Quantitative results for real-world straight-line walking and human-guided trajectory correction experiments. Values in the columns represent the mean and std. dev. across ten trials.
Figure 4: 8, Real world human-guided trajectory correction using our proposed approach. The top row shows a representative interaction sequence. The bottom row reports (a) trunk tilt and (b) planar speed together with the desired and measured headings for the same run. Gray shading marks the interaction, yielding, and replanning interval. Panels (c) and (d) show trunk-tilt and planar-speed histograms across ten trials.
Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy are jointly trained with supervised force and velocity outputs and learned latent context. The estimated force conditions locomotion and additionally generates planar-velocity and yaw-rate commands for leash guidance through an analytical map. In sustained-force simulation sweeps, temporal means of componentwise force root mean square error range from 1.44 to 2.83,N. Compared with a domain-randomized baseline, the framework reduces velocity-tracking and base-orientation error scores by 21.6% and 46.5%, respectively, and increases mean survival from 68.29% to 94.60% in separate sustained-force tests. Unitree Go1 experiments demonstrate stationary vertical and horizontal force estimation, locomotion with an 8.5,kg payload whose weight exceeds the 70,N training force limit, and leash guidance using the same force-estimation interface.
Run Wang, Xu Yang, Alapati Tuerxun +1
Department of Automation, Tsinghua University, Beijing, China
In this paper, we present a robust nonprehensile object transportation framework for quadruped robots. An uncertainty-aware trajectory optimization method generates object motions with minimal closed-loop sensitivity to uncertain parameters. The resulting reference trajectory is tracked using a coupled convex model predictive controller that jointly predicts the CoM dynamics of the quadruped and the payload followed by a whole-body QP that enforces ground reaction constraints. The approach is evaluated through extensive simulations and real-world experiments under variations in the object's inertial parameters. Its performance is compared with fixed-orientation and straight-line trajectories as baseline. The results show that the optimized object motion reduces the sliding by approximately 50% compared with the fixed-orientation baseline and 30% compared with the straight-line baseline, while also achieving lower robot CoM tracking errors.
Ainoor Teimoorzadeh, Riccardo Pretto, Mario Selvaggio +2
Munich Institute of Robotics & Machine Intelligence, Technical University of Munich (TUM), Munich, Germany · Tampere University, Finland · PRISMA Lab, Department of Electrical Engineering and Information Technology, University of Naples Federico II, Via Claudio 21, 80125, Naples, Italy +1
Load transportation with quadruped robots is strongly affected by the dynamics of the physical interface between the robot and the load. Passive spring-based arms reduce weight and complexity compared to active manipulators, but their spring-damper dynamics can introduce oscillatory forces that degrade locomotion stability. This paper derives an extended Zero Moment Point (ZMP) formulation that includes passive payload-interface dynamics, relating stiffness, damping, and payload mass to the stability margin. The analysis shows that underdamped configurations can resonate with locomotion harmonics. Based on this insight, we augment a Single Rigid Body Dynamics model with passive subsystem dynamics and integrate it into a Model Predictive Control framework. In simulation, the proposed controller reduces stability violations by up to 10×, from 7.0% to 0.7%, and increase locomotion efficiency by lowering horizontal ground reaction force effort by up to 15% compared to a nominal baseline. Hardware experiments with a 2kg payload show stable locomotion under pull-release disturbances where the nominal controller fails. The same model also enables end-effector tracking through passive arm dynamics without direct arm actuation.
Giovanni B. Dessy, Lorenzo Amatucci, Victor Barasuol +1
Dynamic Legged Systems Lab, Istituto Italiano di Tecnologia (IIT).