We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Figures & tables
Fig. 1: From a single view of an object, our ZeroBot framework generates a complete object mesh using an image-to-3D model, which is used in simulation for massively parallel reinforcement learning with an efficient action space. The policy is learned in minutes, with no human demonstrations or policy pre-training, and deployed zero-shot on the real robot.
Fig. 2: The three stages of our ZeroBot framework. First, the robot observes an object from a single view. It generates an object mesh from the RGB image using an image-to-3D model, and aligns it with the depth image to recover scale. Second, the generated mesh is used for simulation. The task is specified by providing a goal pose, which defines a dense reward function. Massively parallel RL is used to learn a policy which can manipulate the object into that goal pose. To accelerate training, we use a contact action space which selects a high-value contact state using the learned value function. During deployment, the robot moves into this state using motion planning. Once in contact, it executes actions from its policy.
Fig. 3: An overview of our contact action space on a multi-stage nudge-and-grasp task. (1) First we uniformly sample contact locations from the surface of the generated mesh (a subset of contacts is shown here for visual clarity). Each point shows a sampled gripper pose for nudging (top row) or grasping (bottom row). (2) Then invalid contact states are eliminated, e.g. those colliding with the rest of the object, and the learned value function from RL is used to compute the value of each valid state. Purple states have lower value, and yellow states have higher value. (3) At execution, the max-value state is chosen. (4) The robot moves into contact using motion planning, and then (5) executes policy actions. Meshes shown have been approximated for fast collision checking [ 37 ] . This action space enables policy learning from scratch in minutes.
Fig. 4: The six tasks studied for the end-to-end real world evaluation. Top: MultiStageBook , PushUpRamp , SlideBowl . Bottom: NovelPoses , PullFromShelf , and GraspHandle .
Fig. 7: Qualitative results comparing mesh generation methods. ZeroBot generates a mesh which is complete and more accurate than the baselines, improving sim2real success rates.
Fig. 8: Training times (wall-clock) for each method. Each checkpoint is evaluated in simulation across 256 parallel environments for 10 episodes. Each point on the curve shows the maximum success rate achieved by that run so far. This is averaged across 10 training runs, and the 95% confidence interval for the mean is plotted. Results show that ZeroBot’s action space significantly accelerates training.
Fig. 9: Top: generated mesh and oracle mesh from manual scanning for the articulated object. Bottom: real-to-sim-to-real results, showing cumulative successes against time (left), where the dashed lines indicate the end of mesh creation, and the real robot successfully executing the policy (right).
Robotic data generation is a promising paradigm for scaling robot learning without collecting large-scale real-world data. However, generating geometrically diverse yet physically valid data for contact-rich tasks remains challenging, especially when success depends on precise geometric interfaces. Standard shape augmentation methods often distort task-critical interfaces, resulting in invalid contact relationships, e.g., fit mismatches or interpenetration, rendering downstream interactions infeasible. To address these limitations, we propose a function-preserving Real-to-Sim-to-Real framework that generates synthetic demonstrations from reconstructed assets without teleoperated source trajectories. Our method augments task-relevant object geometries through constraint-guided mesh deformation, together with physically consistent transfer of task poses and collision proxies. Visual domain randomization is further applied during simulation rollouts, enabling robust zero-shot policy deployment without real-world fine-tuning. Extensive experiments in both real-world and simulation settings demonstrate that our method enables robust generalization across unseen object geometries and diverse visual conditions in contact-rich and long-horizon tasks. Our method provides a practical path toward scalable robot learning for contact-rich tasks via shape deformation.
Tianyi Xiang, Xupeng Xie, Jiahang Cao +3
Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China · Institute of Data Science, The University of Hong Kong, Hong Kong SAR, China
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai +15
NVIDIA · Stanford University · The University of Texas at Austin +2
Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.
Kerui Ren, Yingxiang Xu, Kaiwen Song +4
Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Zhejiang University +3