Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training
Organizations: University of Cambridge · Industrial Next
Abstract
Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. We distinguish world grounding, which aligns simulation with the real system, and behavior grounding, which aligns simulated trajectories with human motion. We build a real2sim2real pipeline that varies these axes independently to generate data for co-training. On a dynamic dexterous pick-and-sort task, fully grounded co-training raises success from 52% to 86%; averaged across configurations, world grounding improves success by 18 percentage points and behavior grounding by 10. Deployed policies behave like a mixture of real-derived and simulation-derived policies, imitating real demonstrations in covered states and relying on simulated behavior elsewhere, which we examine through latent-space analysis. Together, these results suggest complementary roles: world grounding lets policies use simulated experience beyond real-data coverage, while behavior grounding matters mainly when world grounding is imperfect. Grounded simulation remains beneficial when co-training foundation models.
Figures & tables
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Factor | Grounded | Ungrounded |
|---|---|---|
| Wrist cameras | Fitted fisheye/mount; blur px | Ideal fisheye; CAD mount; no blur |
| Head/base/belt poses | Registered using robot motion and AprilTags | Tape-measured head; nominal base orientation; supplied cell layout |
| Arm model/control | Factory kinematics; weighted IK, filtering, nullspace; 29 fitted parameters | Nominal kinematics; damped least-squares IK; calculated gains ( ), no armature |
| Hand response | Fitted delay, deadband, and lag | Vendor defaults; direct commands |
| Policy | Nominal | Red light | Fast belt | Novel obj. | Combined | Shifted | Total |
|---|---|---|---|---|---|---|---|
| /24 | /8 | /8 | /9 | /1 | /26 | /50 | |
| 500M policy, trained from scratch | |||||||
| Real 100 | 16 | 2 | 4 | 4 | 0 | 10 | 26 |
| Fully grounded | 23 | 7 | 6 | 6 | 1 | 20 | 43 |
| World grounded | 21 | 3 | 8 | 6 | 0 | 17 | 38 |
| Behavior grounded | 20 | 3 | 6 | 4 | 1 | 14 | 34 |