DepthWorld: 3D World Model for Robot Manipulation
Organizations: Czech Institute of Informatics, Robotics and Cybernetics Czech Technical University in Prague
Abstract
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.
Figures & tables
| All scenes ( ) | PointWorld-converged subset ( ) | ||||
|---|---|---|---|---|---|
| Metric | PointWorld [ 14 ] | Ours | PointWorld [ 14 ] | Ours | |
| EE median (px) | 5.84 | 0.25 | 5.67 | 0.25 | |
| WE median (px) | 13.19 | 1.03 | 12.52 | 1.02 | |
| RD median (mm) | 111.3 | 14.0 | 104.6 | 13.9 | |
| Mask IoU mean | 0.440 | 0.810 | 0.458 | 0.819 | |
| RGB quality | Depth quality | ||||||
|---|---|---|---|---|---|---|---|
| View | Model | PSNR | SSIM | LPIPS | AbsRel | RMSE | |
| External | RGB | 22.63 | 0.8170 | 0.0930 | — | — | — |
| RGB+D | 24.09 | 0.8333 | 0.0889 | 0.0765 | 0.220 | 0.9432 | |
| RGB+D PM-DPT | 24.11 | 0.8333 | 0.0891 | 0.0782 | 0.219 | 0.9434 | |
| Wrist | RGB | 16.98 | 0.5737 | 0.3228 | — | — | — |
| RGB+D | 17.96 | 0.6106 | 0.3137 | 0.2230 | 0.150 | 0.8182 | |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Full model | EE | WE | RD | IoU |
|---|---|---|---|---|
| (px) | (px) | (mm) | (mean) | |
| Full (all factors) | 0.251 | 1.032 | 14.15 | 0.808 |
| Weight tweaks | ||||
| amplify 3D | 0.251 | 1.027 | 14.02 | 0.808 |
| amplify FM | 0.250 | 1.038 | 14.09 | 0.804 |
| Leave-one-out (drop one factor) | ||||
| RGB quality | Depth quality | ||||||
|---|---|---|---|---|---|---|---|
| View | Model | PSNR | SSIM | LPIPS | AbsRel | RMSE | |
| External | Dual-branch | 21.94 | 0.8039 | 0.1010 | 0.0861 | 0.232 | 0.9383 |
| Channel expansion | 22.22 | 0.8017 | 0.1021 | 0.0796 | 0.220 | 0.9414 | |
| Spatial tiling (ours) | 23.66 | 0.8252 | 0.0945 | 0.0854 | 0.235 | 0.9383 | |
| Wrist | Dual-branch | 16.05 | 0.5441 | 0.3554 | 0.2523 | 0.161 | 0.7685 |
| Channel expansion | 16.34 | 0.5593 | 0.3422 | 0.2801 | 0.158 | 0.7930 | |
| External | Wrist | |||||||
|---|---|---|---|---|---|---|---|---|
| rollout | PSNR | LPIPS | AbsRel | PSNR | LPIPS | AbsRel | ||
| RGBD, S 2 M 2 depth (ours) | 24.07 | 0.0877 | 0.074 | 0.946 | 17.68 | 0.3270 | 0.266 | 0.797 |
| RGBD, ZED NEURAL depth | 24.26 | 0.0881 | 0.079 | 0.939 | 17.77 | 0.3260 | 0.254 | 0.794 |
| RGBD, ZED ULTRA depth | 22.11 | 0.0947 | 0.194 | 0.830 | 16.28 | 0.3395 | 0.461 | 0.655 |
| External | Wrist | |||||||
|---|---|---|---|---|---|---|---|---|
| rollout | PSNR | LPIPS | AbsRel | PSNR | LPIPS | AbsRel | ||
| PM-DPT, factory extrinsics | 24.10 | 0.0891 | 0.075 | 0.945 | 18.01 | 0.3131 | 0.233 | 0.817 |
| PM-DPT, refined extrinsics | 24.13 | 0.0890 | 0.075 | 0.945 | 18.03 | 0.3124 | 0.220 | 0.819 |
| Tracked-point , moving | Chamfer, moving | Chamfer, whole scene | ||||
|---|---|---|---|---|---|---|
| Model | 2 s | 8 s | 2 s | 8 s | 2 s | 8 s |
| PointWorld [ 14 ] vs. raw GT | 23 | 56 | 21 | 58 | 8 | 18 |
| DepthWorld vs. VAE-GT | 31 | 46 | 25 | 32 | 8 | 9 |