Humanoid robots could transport payloads substantially heavier than themselves by pulling passive wheeled vehicles instead of carrying the load. This capability, however, creates a coupled locomotion problem: the robot must maintain persistent upper-body contact while adapting to unknown, configuration-dependent forces arising from the payload, vehicle, and terrain. We present a whole-body control framework for humanoid rickshaw pulling that tracks commanded vehicle motion while preserving balance and stable grasps under uncertain load dynamics. During training, a privileged teacher exploits vehicle states, interaction forces, and load properties. Its actions and latent are distilled into a history-conditioned student that implicitly infers coupled dynamics from proprioceptive responses, followed by reinforcement-learning fine-tuning. Comparisons with \emph{No History} and \emph{Only History} baselines show that the resulting policy achieves accurate vehicle tracking while reducing vehicle oscillation, torso tilt, and actuation cost. Behavioral analysis shows that Unitree G1 propels the rickshaw and generates gait-synchronized whole-body reactions that stabilize its lateral and roll motions. Moreover, pulling redistributes joint effort and yields a lower robot-normalized cost-of-transport proxy than unloaded walking over most tested load--speed conditions. On hardware, a single policy performs starting, sustained pulling, turning, and stopping with both rigid payloads and human passengers, handling a loaded rickshaw mass of up to 115~kg without load-specific retuning. These results demonstrate robust heavy-load transportation through coordinated and persistent humanoid--vehicle interaction.
Figures & tables
Fig. 2 : Humanoid–rickshaw system. (a) The physical setup connects G1 to a passive rickshaw through custom grippers. (b) Relative geometry characterizes the coupled configuration. (c) G1 applies pulling forces through the handles, while the wheels support the load and constrain lateral motion.
Fig. 3 : Teacher–student training and deployment. We first train a teacher with privileged load and interaction information, then distill its actions and latent into a student using only onboard observation history. Joint encoder and policy fine-tuning completes training for deployment without privileged inputs.
Term
Expression
Weight
Task tracking
Rickshaw forward speed
exp[−(vr−vrcmd)2/0.2]
2.0
Rickshaw yaw rate
exp[−(ωr−ωrcmd)2/0.1]
3.0
Stability and gait
Torso upright
exp(−∥gtorso,xy∥22/0.2)
0.5
Posture
exp[−291∑j(Δqj/σj)2]
1.0
TABLE I : Whole-body control rewards and weights.
Parameter
Range
Unit
Rickshaw total mass
U(20,60)
kg
Rickshaw CoM offset
U(−0.10,0.10)
m
Rickshaw inertia scale
U(0.8,1.2)
–
Rolling resistance
U(0.01,0.03)
–
Wheel damping
U(0.01,0.03)
Nms/rad
Torso mass offset
U(−2,2)
kg
TABLE II : Domain-randomization ranges.
Fig. 4 : Vehicle tracking and load generalization. Ours and Only History are compared during constant-speed pulling across loads within and beyond the training range.
Metric
Privileged Teacher
Ours
Only History
No History
Fixed Arm
Tracking
Forward RMSE (m/s)
0.278
0.291
0.292
0.381
0.609
Yaw-rate RMSE (rad/s)
0.0450
0.0535
0.0528
0.0809
0.0703
Rickshaw Motion
Pitch-rate RMS (rad/s)
0.0481
0.0530
0.0546
0.0646
0.0775
Forward-accel. RMS (m/s 2 )
0.995
0.960
0.975
1.205
1.154
TABLE III : Policy comparison.
Fig. 5 : Transport cost and contact coordination. (a–c) Load–speed trends, comparison with walking, and joint-effort redistribution characterize the actuation demands of constant-speed pulling. (d–f) Gait-synchronized hand forces and their relation to centroidal motion illustrate the role of handle interaction in whole-body motion regulation.
Fig. 6 : Real-world deployment and sim-to-real comparison. (a) Load scenarios tested with the same policy. (b) Starting, straight pulling, and stopping responses. (c) Robot- and system-normalized CoT. (d,e) Joint angular-velocity and torque RMS comparisons across three speeds for the seated-G1 scenario; labels identify notable discrepancies.