cs.CVSep 28, 2026

RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

Authors: Jin Hyun Kim, Min Young Kim, Soohwan Song, Daekyum Kim

Organizations: School of Mechanical Engineering, Korea University, Seoul, Republic of Korea · College of AI Convergence, Dongguk University, Seoul, Republic of Korea · School of Smart Mobility, Korea University, Seoul, Republic of Korea

Abstract

Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

    Jul 7, 2026Haoyu Zhao, Xingyue Zhao, Siteng Huang +3Robotic ManipulationFour-Dimensional

  2. GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

    May 20, 2026Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8Video World ModelsRobotic Manipulation

  3. Structured 4D Latent Predictive Model for Robot Planning

    Jul 1, 2026Zhiyi Li, Peilin Wu, Xiaoshen Han +2Robot PlanningInverse Dynamics