VioLA: Learning Generalist Humanoid Control Policies from Human Data
Organizations: Vesoma · ETH Zürich · MPI-IS · ELLIS Institute · University of Tuebingen
Abstract
Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.
Figures & tables
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Episodes | Frames | Hours | Targets |
|---|---|---|---|---|
| Humanoid Everyday | 4,064 | 2,962,764 | 16.46 | Body and hands |
| PSI | 835 | 1,431,454 | 7.95 | Body and hands |
| UnifoLM-WBT | 2,854 | 4,454,240 | 24.75 | Body and hands |
| LeVERB | 3,555 | 713,279 | 3.96 | Body only |
| EgoSuite | 20,680 | 131,010,575 | 727.84 | Body and hands |
| Total | 31,988 | 140,572,312 | 780.96 |
| Backbone family | State conditioning | Motion-target scaling | |
|---|---|---|---|
| 50 | 110 values, tokenized | Identity | |
| GR00T N1.7 | 40 | 46 values, state encoder | Identity |
| DiT4DiT | 50 | 32 values, state encoder | Identity |
| Backbone | Episodes | Chunks | Delay (ms) | Replan interval (ms) |
|---|---|---|---|---|
| GR00T | 161 | 3,871 | ||
| 83 | 2,554 | |||
| DiT4DiT | 57 | 1,035 |