RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts
Organizations: School of Mechanical Engineering, Korea University, Seoul, Republic of Korea · College of AI Convergence, Dongguk University, Seoul, Republic of Korea · School of Smart Mobility, Korea University, Seoul, Republic of Korea
Abstract
Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.
Figures & tables
| Method | Source view [0.25ex] | Novel view [0.25ex] | Robot [0.25ex] | Efficiency [0.25ex] | |||||
|---|---|---|---|---|---|---|---|---|---|
| PSNR | LPIPS | AbsRel | PSNR | LPIPS | AbsRel | EPE (mm) | Time (s) | VRAM (GiB) | |
| Recorded observations | |||||||||
| MoVieS † | 19.71 | 0.306 | 0.306 | 13.93 | 0.489 | 0.373 | 22.2 | 1.04 | 4.46 |
| 4DGT | 16.54 | 0.456 | 0.600 | 12.90 | 0.584 | 0.681 | 31.0 | 4.80 | 3.89 |
| SoM † | 14.22 | 0.443 | 0.383 | 12.14 | 0.557 | 0.429 | 32.2 | 1713 | 2.35 |
| Cross camera | Stage2 | Source PSNR | Novel PSNR | Novel AbsRel | Robot EPE (mm) | Robot depth err. (mm) |
|---|---|---|---|---|---|---|
| Recorded Observations | ||||||
| 25.46 | 16.08 | 0.197 | 8.6 | 13.7 | ||
| — | 21.92 | 15.52 | 0.200 | 8.6 | 13.4 | |
| — | 23.71 | 12.93 | 0.392 | 8.8 | 101.6 | |
| — | — | 19.64 | 12.64 | 0.391 | 8.8 | 100.6 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Loss key | Weight | Loss key | Weight |
|---|---|---|---|
| movies_mse | 1.0 | movies_lpips | 0.5 |
| depth_head | 1.0 | track_head | 10.0 |
| stereo | 1.04051 | track_final | 0.105593 |
| fk_head | 0.464340 | fk_final | 0.157137 |
| metric_final | 0.00182939 | metric_right | 0.00484463 |
| coverage | 0.0589678 | coverage_right | 0.0648524 |
| Selection stage | Episodes |
|---|---|
| ILIAD and PennPAL episodes reserved in the internal split | 1,026 |
| Episodes with a prepared clip and stereo references | 804 |
| Eligible episodes represented in the PointWorld test manifest | 498 |
| Fixed-seed evaluation sample | 256 |
| Signal | Ours | MoVieS | 4DGT | SoM |
|---|---|---|---|---|
| Left-eye visual input | 5 times 3 views | 5 / camera | 17 / camera | 17 / camera |
| Calibrated cameras | Yes | Yes | Yes | Yes |
| Joint/gripper states ‡ | Anchoring/FK | No direct input | No direct input | FK masks only |
| Robot mesh | Geometry prior | No | No | Foreground mask |
| Source-time stereo depth | No | Scale only † | No | Scale only † |
| Right-eye RGB/depth maps | Evaluation only | Evaluation only | Evaluation only | Evaluation only |
| Method | Source SSIM | Source RMSE (m) | Novel (%) |
|---|---|---|---|
| Recorded observations | |||
| MoVieS † | 0.713 | 0.333 | 55.1 |
| 4DGT (17f) | 0.585 | 0.414 | 31.1 |
| SoM † | 0.573 | 0.473 | 38.1 |
| Ours (Stage 1) | 0.817 | 0.257 | 77.2 |
| Ours (Stage 1+2) | 0.866 | 0.255 | 77.5 |
| Method | Novel PSNR (dB) | Novel AbsRel | Robot EPE (mm) |
|---|---|---|---|
| Recorded observations | |||
| MoVieS † | 13.93 [13.77, 14.09] | 0.373 [0.356, 0.390] | 22.2 [20.5, 24.2] |
| 4DGT (17f) | 12.90 [12.75, 13.05] | 0.681 [0.659, 0.704] | 31.0 [28.8, 33.2] |
| SoM † | 12.14 [12.00, 12.28] | 0.429 [0.413, 0.445] | 32.2 [30.1, 34.4] |
| Ours (Stage 1) | 15.52 [15.35, 15.70] | 0.200 [0.193, 0.208] | 8.6 [7.8, 9.6] |
| Ours (Stage 1+2) | 16.08 [15.90, 16.27] | 0.197 [0.190, 0.205] | 8.6 [7.7, 9.6] |
| Source | Novel view | ||||
|---|---|---|---|---|---|
| Setting | Frames | PSNR | PSNR | LPIPS | AbsRel |
| Recorded | 5 | 15.63 | 12.78 | 0.603 | 0.803 |
| Recorded | 17 | 16.54 | 12.90 | 0.584 | 0.681 |
| Cosmos 3 | 5 | 15.20 | 12.69 | 0.608 | 0.815 |
| Cosmos 3 | 17 | 15.78 | 12.78 | 0.591 | 0.697 |
| Recorded | Cosmos 3 rollouts | |||
|---|---|---|---|---|
| Method | Source | Novel | Source | Novel |
| 4DGT ‡ | 0.320 | 0.411 | 0.338 | 0.417 |
| SoM †‡ | 0.402 | 0.489 | 0.411 | 0.498 |
| Configuration | Robot EPE (mm) | Robot depth error (mm) |
|---|---|---|
| Full Stage 1 | 8.64 | 13.4 |
| w/o FK conditioning | 30.14 | 21.5 |
| w/o mesh anchoring | 8.80 | 36.5 |
| w/o both | 30.11 | 41.6 |
| Input | Stage | Novel PSNR (dB) | AbsRel | Robot EPE (mm) | Robot depth error (mm) |
|---|---|---|---|---|---|
| Recorded | Stage 1 | 0.00 | -0.004 | -0.16 | +1.7 |
| Recorded | Stage 1+2 | +0.01 | -0.004 | -0.18 | +0.9 |
| Cosmos 3 | Stage 1+2 | +0.01 | -0.004 | -0.20 | +0.9 |
| Updates | Source PSNR | Novel PSNR | Source LPIPS | Source AbsRel | Robot EPE (mm) | Time (s) |
|---|---|---|---|---|---|---|
| 0 | 22.59 | 16.02 | 0.182 | 0.178 | 6.21 | – |
| 1 | 24.94 | 16.39 | 0.169 | 0.177 | 6.22 | 2.61 |
| 2 | 25.59 | 16.44 | 0.168 | 0.177 | 6.22 | 4.78 |
| 3 | 25.80 | 16.47 | 0.171 | 0.176 | 6.22 | 6.95 |
| 4 | 25.87 | 16.48 | 0.173 | 0.176 | 6.22 | 9.12 |
| Input | Novel PSNR (dB) | Novel AbsRel | Robot EPE (mm) | Robot depth error (mm) |
|---|---|---|---|---|
| Recorded | 15.32 | 0.293 | 9.0 | 20.7 |
| Cosmos 3 | 14.98 | 0.298 | 9.0 | 20.7 |
| Input | PSNR (dB) | AbsRel | (%) | Far AbsRel |
|---|---|---|---|---|
| Original RGB | 23.35 | 0.160 | 80.3 | 0.191 |
| down/up RGB | 20.76 | 0.227 | 63.3 | 0.482 |
| Latent adapter | 22.37 | 0.228 | 62.5 | 0.411 |