Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is prohibitively expensive, difficult to scale, and impractical. Generative AI models could remove this bottleneck, but today's models are built for conventional displays and cannot generate the high-resolution, stereoscopic 360∘ content required for immersive viewing. Further, temporal and stereo inconsistencies that may be tolerable on conventional displays can become highly disruptive when viewed through an immersive headset. Here, we address this gap with a zero-shot generative pipeline that extends existing video diffusion models into 4K stereoscopic 360∘ videos. Inspired from binocular vision and depth perception, we develop an epipolar-aware 360∘ image matching metric that captures the temporal and stereo geometric inconsistencies across views. We then use this metric as a preference signal for direct preference optimization with limited training data. Our work enables 360∘ stereo video generation and provides a scalable path for bringing generative content to immersive displays, allowing diverse mixed reality experiences on demand.
Figures & tables
Pano- Wan [ 42 ]
Dissolve- Stereo [ 37 ]
Stereo- World [ 44 ]
Stereo- WM [ 38 ]
Pano- World-X [ 47 ]
Ours
Output domain
360∘ panoramic
✓
✗
✗
✗
✓
✓
Stereo
✗
✓
✓
✓
✗
✓
Geometry
Depth-guided
✗
✓
✓
✓
(✓)
✓
Epipolar constraint
✗
✗
(✓)
(✓)
✗
✓
Table 1: Comparison of related work on panoramic and stereo video generation. Each criterion is fully ✓ , partially (✓) , or not met ✗ . Prior work covers either the panoramic axis or the stereo axis, but not both, and geometric structure enters as depth conditioning rather than as an explicit epipolar constraint.
Figure 1: Illustration of the epipolar constraint. The observation on the left image plane Pl determines a viewing ray, shown by the solid blue line. Because the depth is unknown, the observed 3D point may lie anywhere along this ray, as indicated by the white circles. When these possible 3D locations are projected into the right camera, their projections all lie on the yellow epipolar line LR . Therefore, the point corresponding to the left observation must be found on LR , rather than anywhere in the right image.
Figure 2: Tangent Sampson Error. We compute the error by taking a point from one view and re-projecting it into another view and computing the reprojection error.
Figure 3: Overview. We present a 360 ∘ stereo generation pipeline to convert user provided text or image prompts into headset viewable content. Our pipeline works in three main stages a training-free stereo approach to provide initial stereo estimates, a preference based finetuning approach without requiring expensive 360 ∘ stereo data collection and a final post-processing stage for immersive viewing.
Figure 4: Stage 1: Zero-shot 360 ∘ Stereo Generation. Our training-free stage can produce a large collection of 360 ∘ stereo videos without requiring any training data. We perform depth-wise spherical warping to get the right view from the generated left view and in-paint the occluded regions using the video diffusion model.
Figure 5: Stereo Scorer. Our panoramic stereo scorer metric performs robust feature matching to find compatible point correspondences for epipolar error computation.
Figure 6: DPO. We finetune a pre-trained panoramic video model using low-rank neural adaptors over stereo preference signal obtained from (good, bad) video pairs.
Figure 7: Rigid Stereo to ODS Correction.
Figure 8: Flow-DPO training. The reward margin creeps up while the policy’s divergence from the reference accelerates, so we stop early rather than train to convergence.
Figure 10: Object motion breaks the rigid two-view model. A passenger turns his head and torso between frame t and t+20 while the background barely moves. MEt3R scores every pixel and charges the motion as inconsistency, whereas PEGS votes the moving tracks out (red, which also covers non-rigid foliage) and scores only the static scene: 1.10 vs. 2.54 mrad if they were retained.
Figure 11: DPO refinement. Our DPO refinement recovers fine geometry existing in the source left video lost during diffusion based inpainting over warped right view.
PEGS (mrad)
MEt3R [ 1 ] (unitless)
Axis
Stage 1
Stage 2
Stage 1
Stage 2
Stereo
3.121
2.640
0.096
0.081
Temporal
1.494
1.263
0.061
0.073
Diagonal
2.958
4.375
0.104
0.104
Table 3: Evaluating PEGS and MEt3R on the two generation sets with 500 videos each; lower is better. PEGS is a tangent-space angular residual in milliradians, MEt3R a unitless feature-space dissimilarity, so the two scales are not comparable to each other; only down each column.
Figure 12: Comparing epipolar against DissolveStereo [ 37 ] . We evaluate DissolveStereo’s zero-shot stereo adaptation procedure for panoramic stereo generation and compare against our method in terms of epipolar mismatches. Red indicates more epipolar mismatch and Green indicates less epipolar mismatches.
Figure 13: Camera Controlled Stereo [OmniRoam] vs Ours. Our DPO based stereo adaptation performs better than adapting camera-conditioned video generation models such as OmniRoam which suffers from differing left-right generated views.
Method
Inconsistent
Inconsistent (%)
p95 (mrad)
DissolveStereo
282,949
8.2
15.4
Ours
88,459
2.6
8.4
Table 4: Epipolar consistency against DissolveStereo, over 50 clips sharing the same prompt, seed and left eye, so only the right eye differs. A correspondence is inconsistent when it lies ≥20 mrad off its epipolar curve. Lower is better.
Method
Inconsistent (%)
p95 (mrad)
Disparity (mrad)
Pose-shift
46.0
132.9
46.0
Ours
1.2
8.2
21.5
Table 5: Against a camera pose-shift baseline, over 10 clips sharing the same OmniRoam left eye, so only the right eye differs. Lower is better; disparity is reported to show the comparison is not won by a narrower baseline.
Figure 14: Vertical disparity along one scanline, measured between the delivered eyes. Zero is the ground truth for a fusable pair. The rigid rendering departs from it most where the baseline aligns with the gaze; the omnidirectional conversion holds it at zero throughout.
Figure 15: Met3R vs PEGS. We perform DPO finetune ablation on both PEGS and Met3R with our framework and interestingly observe the tail-end distribution improvement offered by PEGS while Met3R barely shows any improvement.
Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view trajectories, amplifying cross-view inconsistency and temporal drift. We argue that 360° video generation offers a natural solution: panoramic coverage simplifies trajectory design and provides a strong global context for maintaining coherence. We introduce Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion, a controllable 360° video generation framework that synthesizes high-fidelity videos from sparse 360° inputs. The key idea is an explicit 3D Cache, reconstructed from the input, which serves as a geometric scaffold for any user-defined camera path. This allows the diffusion model to focus on photorealistic texture refinement while the 3D Cache enforces global geometric consistency. Experiments show that Pantheon360 achieves superior visual quality and unmatched geometric coherence, enabling reliable and flexible 360° scene generation for downstream simulation and digital-twin applications.
Ting-Hsuan Chen, Ying-Huan Chen, Tao Tu +10
University of Southern California · National Yang Ming Chiao Tung University · Cornell University +1
While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored challenge. We present a pose-free and training-free method that leverages an off-the-shelf monocular video generation model to produce immersive 3D videos. Our approach first warps the generated monocular video into pre-defined camera viewpoints using estimated depth information, then applies a novel \textit{frame matrix} inpainting framework. This framework utilizes the original video generation model to synthesize missing content across different viewpoints and timestamps, ensuring spatial and temporal consistency without requiring additional model fine-tuning. Moreover, we develop a \dualupdate~scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. The resulting multi-view videos are then adapted into stereoscopic pairs or optimized into 4D Gaussians for spatial video synthesis. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, such as Sora, Lumiere, WALT, and Zeroscope. The experiments demonstrate that our method has a significant improvement over previous methods. Project page at: https://daipengwa.github.io/S-2VG_ProjectPage/
Peng Dai, Feitong Tan, Qiangeng Xu +6
Department of Electrical and Electronic Engineering at The University of Hong Kong, Hong Kong.
Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Plücker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.
Jiaming Tan, Zhen Li, Shuwei Shi +6
Alaya Lab · Beijing Institute of Technology · Shanghai Innovation Institute