RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts
Authors: Jin Hyun Kim, Min Young Kim, Soohwan Song, Daekyum Kim
Organizations: School of Mechanical Engineering, Korea University, Seoul, Republic of Korea · College of AI Convergence, Dongguk University, Seoul, Republic of Korea · School of Smart Mobility, Korea University, Seoul, Republic of Korea
Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.
Figures & tables
Figure 1: Overview of RoGSW4RLD for metric 4D lifting of robot-camera rollouts. Given synchronized multi-view videos, RoGSW4RLD reconstructs a shared, time-queryable metric 4D Gaussian scene that supports rendering from unseen viewpoints at arbitrary query times. Compared with independent camera-wise reconstruction using MoVieS ( Lin et al., 2026 ) , followed by stereo-depth metric normalization and calibrated merging in the robot frame, joint lifting substantially improves metric depth, robot-motion accuracy, and novel-view rendering, with Stage 2 further refining the reconstructed field. Bars report results on recorded observations from Table 1 .
Figure 2: Robot-camera setup and limitations of post-hoc field merging. (a) PointWorld-DROID setup with a moving wrist camera and fixed external cameras. (b) Independent reconstruction retains inconsistent geometry after calibrated alignment, while our joint metric reconstruction (Stage 1) forms a more coherent scene. (c) Per-view renderings and merged geometry show reduced duplication and misalignment across cameras.
Figure 3: Overview of RoGSW4RLD. Stage 1 performs kinematically conditioned metric 4D lifting by jointly processing multi-view videos with calibrated camera trajectories, robot-mesh source anchoring, and FK-conditioned learned deformation. Stage 2 recurrently refines the resulting Gaussian field using contribution-weighted rendering feedback, while time-shared corrections preserve each Gaussian’s temporal displacement across query times.
(a) Robot Mesh Anchoring (Where the robot is)
(b) FK Conditioning (How the robot moves)
Method
Source view [0.25ex]
Novel view [0.25ex]
Robot [0.25ex]
Efficiency [0.25ex]
PSNR ↑
LPIPS ↓
AbsRel ↓
PSNR ↑
LPIPS ↓
AbsRel ↓
EPE (mm) ↓
Time (s) ↓
VRAM (GiB) ↓
Recorded observations
MoVieS †
19.71
0.306
0.306
13.93
0.489
0.373
22.2
1.04
4.46
4DGT
16.54
0.456
0.600
12.90
0.584
0.681
31.0
4.80
3.89
SoM †
14.22
0.443
0.383
12.14
0.557
0.429
32.2
1713
2.35
Table 1: Quantitative results on recorded observations and Cosmos 3 rollouts. RoGSW4RLD improves rendering, metric depth, and robot-displacement accuracy over camera-wise baselines. Both input settings are evaluated against recorded references.
Figure 4: Qualitative comparison on action-conditioned Cosmos 3 rollouts. Across ILIAD and PennPAL scenes, RoGSW4RLD produces cleaner reconstructions with less ghosting and sharper robot and object boundaries than the baselines at both input and novel times. The bottom row shows the corresponding generated Cosmos 3 frames, rather than recorded observations; asterisks denote non-input query times for RoGSW4RLD and MoVieS.
Table 3: Stage 1 loss weights. Keys identify separately weighted terms in the training configuration. Values are reported to at most six significant figures.
Selection stage
Episodes
ILIAD and PennPAL episodes reserved in the internal split
1,026
Episodes with a prepared clip and stereo references
804
Eligible episodes represented in the PointWorld test manifest
498
Fixed-seed evaluation sample
256
Appendix
Table 4: Evaluation-cohort construction. Counts refer to episodes, not camera streams or temporal queries.
Signal
Ours
MoVieS
4DGT
SoM
Left-eye visual input
5 times × 3 views
5 / camera
17 / camera
17 / camera
Calibrated cameras
Yes
Yes
Yes
Yes
Joint/gripper states ‡
Anchoring/FK
No direct input
No direct input
FK masks only
Robot mesh
Geometry prior
No
No
Foreground mask
Source-time stereo depth
No
Scale only †
No
Scale only †
Right-eye RGB/depth maps
Evaluation only
Evaluation only
Evaluation only
Evaluation only
Appendix
Table 5: Signals used in the main evaluation. “Scale only” denotes one scalar per camera computed at the five designated source times, not dense stereo supervision during test-time reconstruction. Right-eye maps and reference tracks are not supplied directly to reconstruction.
Figure 5: Novel-view reconstruction from recorded observations. RoGSW4RLD shows less ghosting and clearer scene details at held-out right-eye cameras. The bottom row shows recorded references; asterisks mark non-input times for RoGSW4RLD and MoVieS only.
Method
Source SSIM ↑
Source RMSE (m) ↓
Novel δ1 (%) ↑
Recorded observations
MoVieS †
0.713
0.333
55.1
4DGT (17f)
0.585
0.414
31.1
SoM †
0.573
0.473
38.1
Ours (Stage 1)
0.817
0.257
77.2
Ours (Stage 1+2)
0.866
0.255
77.5
Appendix
Table 6: Complementary reconstruction metrics. Means use 256 clips for MoVieS, 4DGT, and our model, and 246 successful clips for SoM. † : one source-stereo scale scalar per camera before reconstruction.
Method
Novel PSNR (dB) ↑
Novel AbsRel ↓
Robot EPE (mm) ↓
Recorded observations
MoVieS †
13.93 [13.77, 14.09]
0.373 [0.356, 0.390]
22.2 [20.5, 24.2]
4DGT (17f)
12.90 [12.75, 13.05]
0.681 [0.659, 0.704]
31.0 [28.8, 33.2]
SoM †
12.14 [12.00, 12.28]
0.429 [0.413, 0.445]
32.2 [30.1, 34.4]
Ours (Stage 1)
15.52 [15.35, 15.70]
0.200 [0.193, 0.208]
8.6 [7.8, 9.6]
Ours (Stage 1+2)
16.08 [15.90, 16.27]
0.197 [0.190, 0.205]
8.6 [7.7, 9.6]
Appendix
Table 7: Main-comparison means and 95% clip-bootstrap confidence intervals from 10,000 resamples. MoVieS, 4DGT, and our model use 256 clips; SoM uses 246. † : one source-stereo scale scalar per camera.
Source
Novel view
Setting
Frames
PSNR ↑
PSNR ↑
LPIPS ↓
AbsRel ↓
Recorded
5
15.63
12.78
0.603
0.803
Recorded
17
16.54
12.90
0.584
0.681
Cosmos 3
5
15.20
12.69
0.608
0.815
Cosmos 3
17
15.78
12.78
0.591
0.697
Appendix
Table 8: 4DGT with five or 17 input frames per camera under the same camera-wise reconstruction and calibrated-merging protocol on 256 clips.
Recorded
Cosmos 3 rollouts
Method
Source
Novel
Source
Novel
4DGT ‡
0.320
0.411
0.338
0.417
SoM †‡
0.402
0.489
0.411
0.498
Appendix
Table 9: Target-assisted median-scale depth diagnostic (AbsRel ↓ ). † : source-stereo scaling before reconstruction. ‡ : per-image target-assisted scaling used only here, not in the main comparison.
Configuration
Robot EPE (mm) ↓
Robot depth error (mm) ↓
Full Stage 1
8.64
13.4
w/o FK conditioning
30.14
21.5
w/o mesh anchoring
8.80
36.5
w/o both
30.11
41.6
Appendix
Table 10: Robot-specific interventions in Stage 1 on 256 recorded clips. Table 2 reports the corresponding Stage 1+2 interventions for recorded observations and Cosmos 3 rollouts.
Input
Stage
Δ Novel PSNR (dB)
Δ AbsRel
Δ Robot EPE (mm)
Δ Robot depth error (mm)
Recorded
Stage 1
0.00
-0.004
-0.16
+1.7
Recorded
Stage 1+2
+0.01
-0.004
-0.18
+0.9
Cosmos 3
Stage 1+2
+0.01
-0.004
-0.20
+0.9
Appendix
Table 11: Score changes from learned motion to analytic FK transport. Entries are analytic-FK minus learned-motion scores; positive PSNR and negative errors favor analytic transport.
Updates r
Source PSNR ↑
Novel PSNR ↑
Source LPIPS ↓
Source AbsRel ↓
Robot EPE (mm) ↓
Time (s) ↓
0
22.59
16.02
0.182
0.178
6.21
–
1
24.94
16.39
0.169
0.177
6.22
2.61
2
25.59
16.44
0.168
0.177
6.22
4.78
3
25.80
16.47
0.171
0.176
6.22
6.95
4
25.87
16.48
0.173
0.176
6.22
9.12
Appendix
Table 12: Refinement sweep on the first 32 evaluation clips under the fixed-window protocol. Quality metrics use all 32 clips. Refinement time uses CUDA-synchronized measurements on 27 clips after excluding each process’s warm-up clip; lifting and final query rendering are excluded.
Input
Novel PSNR (dB) ↑
Novel AbsRel ↓
Robot EPE (mm) ↓
Robot depth error (mm) ↓
Recorded
15.32
0.293
9.0
20.7
Cosmos 3
14.98
0.298
9.0
20.7
Appendix
Table 13: Optional latent-token interface with Stage 1+2 and four refinement updates on 256 clips. Robot EPE measures track displacement; robot depth error measures source-view placement on the reference robot core. Cosmos 3 rollouts supply generated latents directly.
Input
PSNR (dB) ↑
AbsRel ↓
δ1 (%) ↑
Far AbsRel ↓
Original RGB
23.35
0.160
80.3
0.191
8× down/up RGB
20.76
0.227
63.3
0.482
Latent adapter
22.37
0.228
62.5
0.411
Appendix
Table 14: Stage 1-only spatial-resolution diagnostic on the first 16 evaluation clips under the fixed-window protocol. Metrics use source views; “Far” restricts AbsRel to reference camera- z depths between 1.5 and 5 m. Every input frame and camera is area-downsampled and bilinearly restored in the 8× RGB variant.
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow (RGB-DF) provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to low-level end-effector actions demanded by robotic systems, narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.
Haoyu Zhao, Xingyue Zhao, Siteng Huang +3
DAMO Academy, Alibaba Group · Hong Kong Embodied AI Lab · CUHK +1
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the generated videos appear plausible, yet lack the physical grounding required for reliable action execution, such as robot manipulation. We present GEM-4D, a geometry-grounded video world model that resolves this limitation by injecting dense 4D correspondence supervision distilled from a pretrained geometry foundation model into the video generative backbone during training. This supervision enables the model to jointly capture appearance and geometric structure while retaining a single-stream architecture with no additional inference cost. We further introduce an inverse dynamics module that converts correspondence-consistent video rollouts into executable robot trajectories, enabling direct deployment in both real-world and simulated manipulation. GEM-4D achieves state-of-the-art performance on both video prediction and geometric consistency across both simulation and realistic scenarios and improves real-world manipulation success from 61% to 81%. Additional results are available at https://gem-4d.github.io/.
Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8
Harvard AI and Robotics Lab, Harvard University · Media Lab and EECS, MIT · MIT-IBM Watson AI Lab +1
Video predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning, and flexible decision-making. However, prevailing approaches often operate on 2D video sequences, inherently lacking the 3D geometric understanding necessary for precise spatial reasoning and physical consistency. We introduce a Structured 4D Latent Predictive Model, which predicts the evolution of a scene's 3D structure in a structured latent space conditioned on observations and textual instructions. Our representation encodes the scene holistically and can be decoded into diverse 3D formats, enabling a more complete and 3D consistent scene understanding. This structured 4D latent predictive model serves as a planner, generating future scenes that are translated into executable actions by a goal-conditioned inverse dynamics module. Experiments demonstrate that our model generates futures with strong visual quality, substantially better 3D consistency and multi-view coherence compared to state-of-the-art video-based planners. Consequently, our full planning pipeline achieves superior performance on complex manipulation tasks, exhibits robust generalization to novel visual conditions, and proves effective on real-world robotic platforms. Our website is available at https://structured-4d-model.github.io/.