Humanoid robots could transport payloads substantially heavier than themselves by pulling passive wheeled vehicles instead of carrying the load. This capability, however, creates a coupled locomotion problem: the robot must maintain persistent upper-body contact while adapting to unknown, configuration-dependent forces arising from the payload, vehicle, and terrain. We present a whole-body control framework for humanoid rickshaw pulling that tracks commanded vehicle motion while preserving balance and stable grasps under uncertain load dynamics. During training, a privileged teacher exploits vehicle states, interaction forces, and load properties. Its actions and latent are distilled into a history-conditioned student that implicitly infers coupled dynamics from proprioceptive responses, followed by reinforcement-learning fine-tuning. Comparisons with \emph{No History} and \emph{Only History} baselines show that the resulting policy achieves accurate vehicle tracking while reducing vehicle oscillation, torso tilt, and actuation cost. Behavioral analysis shows that Unitree G1 propels the rickshaw and generates gait-synchronized whole-body reactions that stabilize its lateral and roll motions. Moreover, pulling redistributes joint effort and yields a lower robot-normalized cost-of-transport proxy than unloaded walking over most tested load--speed conditions. On hardware, a single policy performs starting, sustained pulling, turning, and stopping with both rigid payloads and human passengers, handling a loaded rickshaw mass of up to 115~kg without load-specific retuning. These results demonstrate robust heavy-load transportation through coordinated and persistent humanoid--vehicle interaction.
Figures & tables
Fig. 2 : Humanoid–rickshaw system. (a) The physical setup connects G1 to a passive rickshaw through custom grippers. (b) Relative geometry characterizes the coupled configuration. (c) G1 applies pulling forces through the handles, while the wheels support the load and constrain lateral motion.
Fig. 3 : Teacher–student training and deployment. We first train a teacher with privileged load and interaction information, then distill its actions and latent into a student using only onboard observation history. Joint encoder and policy fine-tuning completes training for deployment without privileged inputs.
Term
Expression
Weight
Task tracking
Rickshaw forward speed
exp[−(vr−vrcmd)2/0.2]
2.0
Rickshaw yaw rate
exp[−(ωr−ωrcmd)2/0.1]
3.0
Stability and gait
Torso upright
exp(−∥gtorso,xy∥22/0.2)
0.5
Posture
exp[−291∑j(Δqj/σj)2]
1.0
TABLE I : Whole-body control rewards and weights.
Parameter
Range
Unit
Rickshaw total mass
U(20,60)
kg
Rickshaw CoM offset
U(−0.10,0.10)
m
Rickshaw inertia scale
U(0.8,1.2)
–
Rolling resistance
U(0.01,0.03)
–
Wheel damping
U(0.01,0.03)
Nms/rad
Torso mass offset
U(−2,2)
kg
TABLE II : Domain-randomization ranges.
Fig. 4 : Vehicle tracking and load generalization. Ours and Only History are compared during constant-speed pulling across loads within and beyond the training range.
Metric
Privileged Teacher
Ours
Only History
No History
Fixed Arm
Tracking
Forward RMSE (m/s)
0.278
0.291
0.292
0.381
0.609
Yaw-rate RMSE (rad/s)
0.0450
0.0535
0.0528
0.0809
0.0703
Rickshaw Motion
Pitch-rate RMS (rad/s)
0.0481
0.0530
0.0546
0.0646
0.0775
Forward-accel. RMS (m/s 2 )
0.995
0.960
0.975
1.205
1.154
TABLE III : Policy comparison.
Fig. 5 : Transport cost and contact coordination. (a–c) Load–speed trends, comparison with walking, and joint-effort redistribution characterize the actuation demands of constant-speed pulling. (d–f) Gait-synchronized hand forces and their relation to centroidal motion illustrate the role of handle interaction in whole-body motion regulation.
Fig. 6 : Real-world deployment and sim-to-real comparison. (a) Load scenarios tested with the same policy. (b) Starting, straight pulling, and stopping responses. (c) Robot- and system-normalized CoT. (d,e) Joint angular-velocity and torque RMS comparisons across three speeds for the seated-G1 scenario; labels identify notable discrepancies.
Humanoid loco-manipulation of large, heavy objects demands forceful interaction across the entire body. However, such payloads shift a humanoid's center of mass and impose sustained loads across the upper body, challenging balance and command tracking. We present HULK, a whole-body control framework for forceful loco-manipulation. Using model predictive control (MPC) to guide reinforcement learning with predictions of the loaded dynamics, we train two teachers: one tracks arm motions under wrist forces, and the other locomotes while holding large objects against the body. A capture-point control barrier function augments the wrist-force teacher during training to improve balance under load. We distill both teachers into a single policy. Evaluation spans simulation and the Unitree G1. In simulation, the teacher with the barrier function achieves the lowest forward and lateral velocity tracking errors at 10 kg per arm among evaluated controllers and reduces aggregate divergent component of motion (DCM) excursion magnitude by 35.7% relative to MPC-guided reinforcement learning alone. Our wrist-force teacher withstands torso push disturbances of up to 130 N.
An Dang, Arturo Flores Alvarez, Yu-Ming Chen +5
Amazon · University of Michigan · University of California, Los Angeles +1
Quadruped robots are increasingly expected to carry objects while moving through human environments. But what happens when a person interacts directly with the payload rather than with the robot? If the payload is unrestrained, the robot must distinguish intentional external interactions from ordinary payload motion, while still keeping the load balanced and maintaining stable locomotion. How can a quadruped infer and compliantly respond to such interactions using only onboard measurements? In this work, we develop a force-aware locomotion framework that treats payload interactions as commands that shape the motion of the combined robot-payload system. Our approach separates the learning of force-aware locomotion and force estimation on an unrestrained payload. We combine a compliant load-carrying policy with a causal force estimator, trained through estimator-in-the-loop data aggregation and finetuning, to predict interactions from onboard robot measurements. Our simulations and real-world experiments show that the resulting controller can maintain stable payload-carrying locomotion, yield compliantly to external interactions, and use the inferred force to support human-guided changes in the robot's trajectory.
Shaunak A. Mehta, Mayank Mishra, Prajit KrisshnaKumar +2
Fujitsu Research of America, Pittsburgh, PA, 15213.
Whole-body compliant control is essential for deploying heavy humanoids under high payload in human-centric environments. Most prior force-aware learning-based pipelines focus on end-effector resistance, per-link upper-body springs, or end-effector stiffness modulation, leaving arbitrary-site perturbations on heavy platforms with lower-body engagement largely unaddressed. We close this gap with CompliantWBC comprising: (1) A base policy trained with RL to maximize compliance-fidelity reward, guided by a multi-site whole-body impedance reference controller, extending classical Cartesian impedance to any controlled link; (2) A bounded residual policy that edits the per-link impedance equilibrium over a frozen base, correcting the coarse but structured wrench estimate supplied by a force encoder co-trained behind a gradient barrier; (3) A Phong-weighted force-origin sampler with an axis-decoupled pelvis anchor induces lower-body-inclusive compliance curriculum training via two interpretable parameters. We evaluate CompliantWBC in simulation against both compliant and stiff baselines, achieving best compliant fidelity of 2.58cm deviation from analytical solutions, and demonstrate it on a real heavy humanoid across static/dynamic force reaction, board wiping, squat under payload, and cooperative payload transport. Project website: https://dotandung.github.io/compliantwbc/
Tan-Dzung Do, Cuc T. Trinh, Tuan Dat Phuong +4
VinRobotics · National University of Singapore · VinUniversity +1