Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Figures & tables
Figure 1: PRISM leverages video-to-video (V2V) generation to expand a few real videos into diverse counterfactual interactions—interactions that did not occur in the source videos but could have occurred with different objects. Reconstruction and retargeting yield physically plausible robot–object trajectories for training a unified depth-based humanoid policy that picks up, carries, and drops diverse objects zero-shot in real world deployment. Project website: prism-real2sim2real.github.io .
Figure 2: PRISM Real-to-sim Overview. PRISM turns counterfactual human–object videos into deployable humanoid loco-manipulation skills. It reconstructs the camera, human motion, object geometry, and 6D object motion in a shared world frame, using contact points to constrain object pose optimization under monocular ambiguity. The same anchors are used in retargeting to preserve interaction phases and match robot end-effectors to intended contact points . The retargeted demonstrations train a privileged co-tracking teacher, which is distilled into a depth-based student policy conditioned on onboard depth and joystick commands for zero-shot sim-to-real deployment.
Figure 3: Policy rollout under real-world object variations. Our single unified policy zero-shot transfers to the real robot across object-pose, intra-category, scale, and approach-distance variations. The policy is agnostic to object placement and initial pose ( top-left ), handles objects with different topology and appearance within the same category ( top-right ), and generalizes to large scale variations ( bottom-left ). It further adapts to object distance using onboard depth, either approaching (up to 4.3 feet!) before grasping or directly grasping when the object is within reach ( bottom-right ).
Training Data
OMOMO Train
OMOMO Test
PRISM ID
PRISM OOD
OMOMO (90)
94.44%
91.67%
23.75%
12.50%
PRISM ID(80)
98.89%
100%
96.25%
72.92%
Table 1: Cross-domain evaluation. OMOMO-trained policies perform well on held-out OMOMO interactions but transfer poorly to PRISM. PRISM-trained policies transfer to OMOMO and unseen PRISM-OOD categories; PRISM-ID evaluates the 80 training interactions.
In-domain
Out-of-domain
Metric
Ball
Bin
Barrel
Box
Paper towel
Helmet
Table
Backpack
Lamp
Chair
Kettle
Chick toy
Success / Trials
12 / 15
14 / 15
14 / 15
15 / 15
5 / 5
3 / 5
5 / 5
5 / 5
3 / 5
4 / 5
4 / 5
3 / 5
Success Rate
80%
93%
93%
100%
100%
60%
100%
100%
60%
80%
80%
60%
Table 2: Real-world evaluation. We evaluate our policy on real-world objects. Each in-domain category contains 3 different objects (see Fig. 7 for details). Each object is tested with 5 trials.
Contact-anchored Pose
Contact-anchored Retargeting
Contact Reward
Suc. (ID)
Suc. (OOD)
Baseline
✗
✗
✗
22.50%
12.50%
+ anchored object pose
✓
✗
✗
58.75%
29.17%
+ contact-aware retargeting
✓
✓
✗
86.25%
68.75%
+ contact reward (PRISM)
✓
✓
✓
96.25%
72.92%
Table 3: Ablation study. We progressively add contact-anchored pose reconstruction, retargeting, and contact rewards. Each component improves task success on PRISM-ID and PRISM-OOD.
Figure 4: Initial object-pose coverage. Object poses relative to G1 in 137 generated clips and four upright seed videos. Left: planar positions of generated clips (circles) and seeds (squares). Right: pooled yaw counts in six 30∘ bins modulo 180∘ . Colors distinguish upright and lying-down objects.
Figure 5: Zero-shot elevated pick-up. Though trained only on flat terrain, our policy picks up a box on a 35∘ ramp ( left ) and from a 0.43m elevated support ( right ), without additional policy tuning.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Real-world test objects. In-domain objects are unseen instances of boxes, bins, barrels, and balls. Out-of-domain objects belong to categories absent from the generated training videos. Both groups vary in appearance, geometry, weight, and scale.
Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scale because each collection depends on physical setups, instrumented actors, and robot operation. We present GRAIL, a digital generation pipeline that remains fully virtual until deployment: it composes 3D assets, simulator-ready scenes, and priors from video foundation models (VFMs) to synthesize interactions without rebuilding physical environments or teleoperating the robot. Rather than reconstructing unconstrained in-the-wild videos, GRAIL starts from fully specified 3D configurations in which object geometry, camera parameters, metric scale, environment depth, and a robot-proportioned character are known before video generation and reused during reconstruction. This privileged setup better conditions 4D recovery, allowing model-based object tracking, human motion estimation, and interaction-aware optimization to reconstruct metric 4D human-object interaction (HOI) trajectories with reduced depth ambiguity and morphology mismatch. We retarget the recovered motions to a humanoid robot and train complementary task-general trackers: an object-aware latent adaptor for manipulation and a scene-aware tracker for terrain traversal. GRAIL produces over 20,000 sequences spanning pick-up, object manipulation, sitting, and terrain traversal. Using only GRAIL-generated data, we train egocentric visual policies through a sim-to-real pipeline and deploy them on a Unitree G1 humanoid, achieving 84% real-world success on diverse object pick-up and 90% success on stair-climbing.
Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping.
Recent advances in large-scale pretrained vision-language-action models have improved robot policy learning, but directly deploying such policies in user-specific environments remains challenging due to limited generalization, which inevitably requires collecting a dataset tailored to the target environment. Teleoperation yields well-aligned data but is costly and difficult to scale, whereas simulation scales easily but struggles to resemble the target environment and generate task-specific trajectories. To meet both simultaneously, we propose PRISM, an end-to-end pipeline that generates personalized robotic datasets from a single image and a natural-language instruction. PRISM constructs digital cousin scenes that are semantically and geometrically aligned with the user environment yet diverse at the instance level, and synthesizes executable demonstrations without human teleoperation. Extensive experiments show that policies trained on PRISM-generated datasets outperform those trained on baseline-generated datasets on LIBERO and LIBERO-Plus, achieve up to 100% success rate on three real-world manipulation tasks, and maintain stronger performance when evaluated in environments that differ from those seen during training.