Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. We distinguish world grounding, which aligns simulation with the real system, and behavior grounding, which aligns simulated trajectories with human motion. We build a real2sim2real pipeline that varies these axes independently to generate data for co-training. On a dynamic dexterous pick-and-sort task, fully grounded co-training raises success from 52% to 86%; averaged across configurations, world grounding improves success by 18 percentage points and behavior grounding by 10. Deployed policies behave like a mixture of real-derived and simulation-derived policies, imitating real demonstrations in covered states and relying on simulated behavior elsewhere, which we examine through latent-space analysis. Together, these results suggest complementary roles: world grounding lets policies use simulated experience beyond real-data coverage, while behavior grounding matters mainly when world grounding is imperfect. Grounded simulation remains beneficial when co-training foundation models.
Figures & tables
Figure 1: Overview of world and behavior grounding in real2sim2real co-training. (Left) We vary simulated data along two axes, world grounding and behavior grounding, giving four simulated datasets. (Center) Each is combined with 100 real teleoperation demonstrations to co-train policies from scratch or post-train foundation models. (Right) Hypothesized mechanism (schematic, not real data): a deployed policy imitates the real demonstrations in states they cover, relies on simulated behavior elsewhere, and returns to real-like states once it can.
Figure 2: Grounded co-training improves task success, for both (a) flow-matching policies and (b) foundation models post-trained with sim-real mixture data. All co-trained policies (Ungrounded, Behavior Grounded, World Grounded, Grounded) use 100 real demonstrations plus about 1,500 simulated trajectories. Error bars show 95% Wilson score intervals over 50 real-world trials.
Figure 3: Visual grounding and rollout behavior in observation-embedding PCA projections. (a) CRAFT-style visual transfer places simulated observations closer to real data than native RTX rendering in this projection. (b) A selected rollout of a diagnostic policy trained with 10 real demonstrations and fully ungrounded simulation data moves toward the simulated data region while exhibiting a two-finger pinch, a characteristic of ungrounded simulated trajectories.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Factor
Grounded
Ungrounded
Wrist cameras
Fitted fisheye/mount; blur σ=0.53 px
Ideal 160∘ fisheye; CAD mount; no blur
Head/base/belt poses
Registered using robot motion and AprilTags
Tape-measured head; nominal base orientation; supplied cell layout
In recent years, reinforcement learning (RL) has shown remarkable success in robotics when a fast and accurate simulator is available for a given task. When using RL and simulation, more simulator realism is generally beneficial but becomes harder to obtain as robots are deployed in increasingly complex and widescale domains. In such settings, simulators will likely fail to model all relevant details of a given target task and this observation motivates the study of sim2real with simulators that leave out key task details. In this paper, we formalize and study the abstract sim2real problem: given an abstract simulator that models a target task at a coarse level of abstraction, how can we train a policy with RL in the abstract simulator and successfully transfer it to the real-world? Our first contribution is to formalize this problem using the language of state abstraction from the RL literature. This framing shows that an abstract simulator can be grounded to match the target task if the grounded abstract dynamics take the history of states into account. Based on the formalism, we then introduce a method that uses real-world task data to correct the dynamics of the abstract simulator. We then show that this method enables successful policy transfer both in sim2sim and sim2real evaluation.
Yunfu Deng, Yuhao Li, Josiah P. Hanna
Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI 53706 USA · Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Amherst, MA 01003 USA
Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet world-action models have not previously been shown to transfer from simulation to real robotic manipulation. We study whether a world-action model can be trained from synthetic priors and deployed zero-shot in the real world. To this end, we build upon Cosmos Policy, a video diffusion model adapted for visuomotor control. We construct simulation environments with extensive domain randomization and generate demonstrations using the AnyTask motion planning pipeline. We evaluate our approach across object lifting, drawer opening, and pick-and-place tasks using ∼800 synthetic demonstrations per task and no real demonstrations. When deployed zero-shot on a Franka Robot, our policy attains a 35% average success rate. To our knowledge, this represents the first successful sim-to-real transfer of a world-action model for robotic manipulation.
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai +15
NVIDIA · Stanford University · The University of Texas at Austin +2