Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data
Organizations: University of California, Berkeley
Abstract
Humanoid robots still struggle to plan contact-rich whole-body manipulation from egocentric RGB and proprioception. Workhorse learns such manipulation from robot-free human demonstrations. A visual planner predicts five-link targets: the poses of the torso, both wrists, and both feet. A reinforcement-learning whole-body tracker follows them on the robot. Both policies train separately on the same recorded human poses, without retargeting. We augment the training data of each policy to imitate the errors that the other makes at deployment. On a real Unitree G1, Workhorse sorts boxes with its hands and a kick, catches a thrown box, and topples and climbs a suitcase. During box sorting, we show recoveries after a person pushes the robot or takes the box away. In a simulated copy of the demonstration room, the system completes box sorting in 77% of episodes, and in 64% under 40 N.s pushes. With both policies retrained from the same demonstrations, a simulated second humanoid completes box sorting in 83% of episodes without pushes.
Figures & tables
| Kernel | Frame | Width |
| Style [ 1 ] | ||
| Pelvis attitude | world | |
| Link pose, all 14 | pelvis , yaw | , |
| Link velocity, all 14 | world | , |
| Task | ||
| Torso, wide | world | , |
| Without pushes | With pushes | |||||||
| 1 | 2 | Success | Fall | 1 | 2 | Success | Fall | |
| Ours (full system) | 91 | 84 | 77 | 1 | 86 | 68 | 64 | 15 |
| Joint command [ 7 ] | 73 | 58 | 54 | 1 | 86 | 50 | 42 | 4 |
| Link-local action frame | 34 | 9 | 5 | 77 | 21 | 5 | 4 | 91 |
| w/o projected gravity | 24 | 5 | 1 | 97 | 13 | 2 | 1 | 99 |
| w/o planner augmentation | 36 | 25 | 20 | 79 | 7 | 1 | 0 | 95 |