Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation
Organizations: Kaedim
Abstract
Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however, its value depends on how closely its outcomes track the real robot's. We test whether our reconstruction pipeline, combining metrically scaled object geometry, authored physical parameters and scene reconstruction, reduces disagreement between simulated and real robot scores relative to a default open-source recipe. We constructed two simulated versions of one bimanual robot cell: an authored reconstruction, using object geometry at estimated metric scale, projected textures, authored physics and our own scene splat; and a baseline, referred to as the default reconstruction, using the open-source recipe of a generative single-image mesh, engine-default physics and a Gaussian-splat scene. Both reconstructions use the same object photographs and scene video. Two policies ran five tasks each, giving ten task-policy pairs, which we call cells; each cell was run twenty times in each reconstruction with all other settings held fixed. Both reconstructions were scored against the same real trials, graded by a third-party evaluator. Pearson correlation between the ten simulated and real cell means is r = 0.90 for the authored reconstruction and 0.51 for the default. Mean score error is 6.97 percentage points for the authored reconstruction and 17.54 for the default, a reduction of 10.56 percentage points. These results show that improving the quality of the environment reconstruction through higher visual fidelity, authored physics and metric scale makes the simulation more faithful to the real world and narrows the sim-to-real gap. We release the harness, the per-trial scores, every reported run's configuration, and the assets and scenes of both reconstructions.
Figures & tables
| authored reconstruction | default reconstruction | |
|---|---|---|
| object geometry | mesh reconstructed with our pipeline from a photograph | mesh generated by TRELLIS from the same photograph |
| texture | projected from the same photograph | generated by TRELLIS with the mesh |
| metric scale | estimated by our pipeline from a printed ChArUco board of known size captured with the object | typed in by hand to match the object’s apparent size in the same photograph; no real measurement |
| collision geometry | signed distance field (bowl, cup, can, spray bottle, cube, mug, plate); authored convex hulls (bottles, blocks, cutting board); bounding sphere (baseball); convex decomposition (grapes, bin) | convex decomposition |
| mass, friction, restitution | mass and restitution estimated by a vision-language model per object; friction from a per-class table; material class from the same model | not authored; PhysX defaults, density 1000 kg/m 3 , friction 0.5, restitution 0 |
| scene reconstruction, from the same handheld video | 3D Gaussian splat [ 14 ] trained from the handheld video, aligned by a table-plane fit, cut above the table | 2D Gaussian splat [ 16 ] by a PolaRiS-inspired recipe from the same video, same alignment and cut |
| setting | value |
|---|---|
| control | 30 Hz; absolute joint targets clipped to each joint’s range, as on the evaluator’s controller |
| action chunk | 16 steps per inference ( 0.5), 30 (MolmoAct2) |
| MolmoAct2 solver steps | 10 |
| step budget per trial | bottles 6400, bowls 3600, blocks 3600, latte 3600, clear table 12800, matching the evaluator |
| top camera | rendered 640x480, served at 640x360; 70 degrees below horizontal; 0.72 m ( 0.5) or 0.88 m (MolmoAct2); measured D435 intrinsics per rig |
| object placement | one layout per cell refit from the evaluator’s start frame; up to 1 cm seeded jitter in x and y per trial |
| task (max) | task string | stage | points |
|---|---|---|---|
| bottles in bin (20) | “put 2 glass bottles in the blue bin” | grasped and lifted | 5 |
| in bin, on its side | 8 | ||
| in bin, upright | 10 | ||
| total score of two bottles in bin | 20 | ||
| stack blocks (30) | “stack the red block on the blue block” | attempt | 5 |
| red lifted | 10 |
| measure | authored | default | paired difference, 95% CI | excludes zero |
|---|---|---|---|---|
| score error (percentage points) | 6.97 | 17.54 | +10.56 [+5.03, +17.11] | yes |
| progress disagreement (percentage points) | 15.90 | 24.23 | +8.32 [+1.78, +16.74] | yes |
| Pearson correlation (r) | 0.90 | 0.51 | +0.39 [-0.08, +0.87] | no |
| failure-stage disagreement (percentage points) | 42.00 | 48.50 | +6.50 [-3.50, +17.00] | no |
| task | policy | real | real 95% CI | authored | authored consistent | default | default consistent | margin |
|---|---|---|---|---|---|---|---|---|
| bottles in bin (20) | MolmoAct2 | 9.50 | 7.15 to 11.75 | 5.95 | yes | 6.25 | yes | -0.30 |
| bottles in bin (20) | 0.5 | 10.05 | 7.65 to 12.35 | 9.40 | yes | 8.90 | yes | 0.50 |
| stack bowls (50) | MolmoAct2 | 29.75 | 23.50 to 35.75 | 25.75 | yes | 16.75 | no | 9.00 |
| stack bowls (50) | 0.5 | 29.25 | 23.00 to 35.50 | 30.50 | yes | 23.00 | yes | 5.00 |
| stack blocks (30) | MolmoAct2 | 3.50 | 1.50 to 5.50 | 7.75 | no | 12.75 | no | 5.00 |
| stack blocks (30) | 0.5 | 6.00 | 4.00 to 8.25 | 5.75 | yes | 5.50 | yes | 0.25 |
| cell | authored | default | stages |
|---|---|---|---|
| bottles, MolmoAct2 | 22.0 | 20.0 | 5 |
| bottles, 0.5 | 15.0 | 17.5 | 4 |
| bowls, MolmoAct2 | 7.5 | 25.8 | 6 |
| bowls, 0.5 | 6.7 | 10.0 | 6 |
| blocks, MolmoAct2 | 20.0 | 60.0 | 1 |
| blocks, 0.5 | 13.3 | 15.0 | 3 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| cell | authored, all gate-passing batches | default, all gate-passing batches |
|---|---|---|
| bottles, MolmoAct2 | 3.45 (V2_K_bottles_molmo), REPORTED 5.95 (K2_bot_molmo_plain), 6.50 (A1_bot_molmo_plain) | REPORTED 6.25 (S1_PS_bot_molmo) |
| bottles, 0.5 | REPORTED 9.40 (K1_bot_pi05_plain), 10.80 (FINAL_K_bottles_pi05), 10.85 (A3_bot_pi05_plain) | REPORTED 8.90 (B0_PS_bot_pi05_r2), 9.50 (S6_PS_bot_pi05), 9.65 (A0_PS_bot_pi05_rerun) |
| bowls, MolmoAct2 | 19.75 (six jitter-sweep shards pooled: 04_NS10, JIT25, JIT40, DBG_JIT25, DBG_JIT40, PROBE_NS10; the gate cannot see the placement shift, so they pass it; listed, not reported), 20.00 (N20_K_bowls_molmo), 23.25 (V2_K_bowls_molmo), REPORTED 25.75 (RB_BOWLS_a + RB_BOWLS_b) | REPORTED 16.75 (S2_PS_bowls_molmo) |
| bowls, 0.5 | 28.50 (FINAL_K_bowls_pi05), REPORTED 30.50 (C2_bowls_pi05) | REPORTED 23.00 (S7_PS_bowls_pi05) |
| blocks, MolmoAct2 | REPORTED 7.75 (C1_blocks_molmo), 8.00 (N20_K_blocks_molmo), 9.25 (V2_K_blocks_molmo) | 10.50 (N20_P_blocks_molmo), 12.50 (V2_P_blocks_molmo), REPORTED 12.75 (S3_PS_blocks_molmo) |
| blocks, 0.5 | REPORTED 5.75 (T1_blocks_pi05), 6.00 (FINAL_K_blocks_pi05) | REPORTED 5.50 (ZZ1_PS_blocks_pi05) |