Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
Organizations: Amazon FAR · UC Berkeley · Stanford · Carnegie Mellon University
Abstract
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Figures & tables
| Training Data | OMOMO Train | OMOMO Test | PRISM ID | PRISM OOD |
|---|---|---|---|---|
| OMOMO (90) | 94.44% | 91.67% | 23.75% | 12.50% |
| PRISM ID(80) | 98.89% | 100% | 96.25% | 72.92% |
| In-domain | Out-of-domain | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | Ball | Bin | Barrel | Box | Paper towel | Helmet | Table | Backpack | Lamp | Chair | Kettle | Chick toy |
| Success / Trials | 12 / 15 | 14 / 15 | 14 / 15 | 15 / 15 | 5 / 5 | 3 / 5 | 5 / 5 | 5 / 5 | 3 / 5 | 4 / 5 | 4 / 5 | 3 / 5 |
| Success Rate | 80% | 93% | 93% | 100% | 100% | 60% | 100% | 100% | 60% | 80% | 80% | 60% |
| Contact-anchored Pose | Contact-anchored Retargeting | Contact Reward | Suc. (ID) | Suc. (OOD) | |
|---|---|---|---|---|---|
| Baseline | ✗ | ✗ | ✗ | 22.50% | 12.50% |
| + anchored object pose | ✓ | ✗ | ✗ | 58.75% | 29.17% |
| + contact-aware retargeting | ✓ | ✓ | ✗ | 86.25% | 68.75% |
| + contact reward (PRISM) | ✓ | ✓ | ✓ | 96.25% | 72.92% |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.