Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is prohibitively expensive, difficult to scale, and impractical. Generative AI models could remove this bottleneck, but today's models are built for conventional displays and cannot generate the high-resolution, stereoscopic 360∘ content required for immersive viewing. Further, temporal and stereo inconsistencies that may be tolerable on conventional displays can become highly disruptive when viewed through an immersive headset. Here, we address this gap with a zero-shot generative pipeline that extends existing video diffusion models into 4K stereoscopic 360∘ videos. Inspired from binocular vision and depth perception, we develop an epipolar-aware 360∘ image matching metric that captures the temporal and stereo geometric inconsistencies across views. We then use this metric as a preference signal for direct preference optimization with limited training data. Our work enables 360∘ stereo video generation and provides a scalable path for bringing generative content to immersive displays, allowing diverse mixed reality experiences on demand.
Figures & tables
Pano- Wan [ 42 ]
Dissolve- Stereo [ 37 ]
Stereo- World [ 44 ]
Stereo- WM [ 38 ]
Pano- World-X [ 47 ]
Ours
Output domain
360∘ panoramic
✓
✗
✗
✗
✓
✓
Stereo
✗
✓
✓
✓
✗
✓
Geometry
Depth-guided
✗
✓
✓
✓
(✓)
✓
Epipolar constraint
✗
✗
(✓)
(✓)
✗
✓
Table 1: Comparison of related work on panoramic and stereo video generation. Each criterion is fully ✓ , partially (✓) , or not met ✗ . Prior work covers either the panoramic axis or the stereo axis, but not both, and geometric structure enters as depth conditioning rather than as an explicit epipolar constraint.
Figure 1: Illustration of the epipolar constraint. The observation on the left image plane Pl determines a viewing ray, shown by the solid blue line. Because the depth is unknown, the observed 3D point may lie anywhere along this ray, as indicated by the white circles. When these possible 3D locations are projected into the right camera, their projections all lie on the yellow epipolar line LR . Therefore, the point corresponding to the left observation must be found on LR , rather than anywhere in the right image.
Figure 2: Tangent Sampson Error. We compute the error by taking a point from one view and re-projecting it into another view and computing the reprojection error.
Figure 3: Overview. We present a 360 ∘ stereo generation pipeline to convert user provided text or image prompts into headset viewable content. Our pipeline works in three main stages a training-free stereo approach to provide initial stereo estimates, a preference based finetuning approach without requiring expensive 360 ∘ stereo data collection and a final post-processing stage for immersive viewing.
Figure 4: Stage 1: Zero-shot 360 ∘ Stereo Generation. Our training-free stage can produce a large collection of 360 ∘ stereo videos without requiring any training data. We perform depth-wise spherical warping to get the right view from the generated left view and in-paint the occluded regions using the video diffusion model.
Figure 5: Stereo Scorer. Our panoramic stereo scorer metric performs robust feature matching to find compatible point correspondences for epipolar error computation.
Figure 6: DPO. We finetune a pre-trained panoramic video model using low-rank neural adaptors over stereo preference signal obtained from (good, bad) video pairs.
Figure 7: Rigid Stereo to ODS Correction.
Figure 8: Flow-DPO training. The reward margin creeps up while the policy’s divergence from the reference accelerates, so we stop early rather than train to convergence.
Figure 10: Object motion breaks the rigid two-view model. A passenger turns his head and torso between frame t and t+20 while the background barely moves. MEt3R scores every pixel and charges the motion as inconsistency, whereas PEGS votes the moving tracks out (red, which also covers non-rigid foliage) and scores only the static scene: 1.10 vs. 2.54 mrad if they were retained.
Figure 11: DPO refinement. Our DPO refinement recovers fine geometry existing in the source left video lost during diffusion based inpainting over warped right view.
PEGS (mrad)
MEt3R [ 1 ] (unitless)
Axis
Stage 1
Stage 2
Stage 1
Stage 2
Stereo
3.121
2.640
0.096
0.081
Temporal
1.494
1.263
0.061
0.073
Diagonal
2.958
4.375
0.104
0.104
Table 3: Evaluating PEGS and MEt3R on the two generation sets with 500 videos each; lower is better. PEGS is a tangent-space angular residual in milliradians, MEt3R a unitless feature-space dissimilarity, so the two scales are not comparable to each other; only down each column.
Figure 12: Comparing epipolar against DissolveStereo [ 37 ] . We evaluate DissolveStereo’s zero-shot stereo adaptation procedure for panoramic stereo generation and compare against our method in terms of epipolar mismatches. Red indicates more epipolar mismatch and Green indicates less epipolar mismatches.
Figure 13: Camera Controlled Stereo [OmniRoam] vs Ours. Our DPO based stereo adaptation performs better than adapting camera-conditioned video generation models such as OmniRoam which suffers from differing left-right generated views.
Method
Inconsistent
Inconsistent (%)
p95 (mrad)
DissolveStereo
282,949
8.2
15.4
Ours
88,459
2.6
8.4
Table 4: Epipolar consistency against DissolveStereo, over 50 clips sharing the same prompt, seed and left eye, so only the right eye differs. A correspondence is inconsistent when it lies ≥20 mrad off its epipolar curve. Lower is better.
Method
Inconsistent (%)
p95 (mrad)
Disparity (mrad)
Pose-shift
46.0
132.9
46.0
Ours
1.2
8.2
21.5
Table 5: Against a camera pose-shift baseline, over 10 clips sharing the same OmniRoam left eye, so only the right eye differs. Lower is better; disparity is reported to show the comparison is not won by a narrower baseline.
Figure 14: Vertical disparity along one scanline, measured between the delivered eyes. Zero is the ground truth for a fusable pair. The rigid rendering departs from it most where the baseline aligns with the gaze; the omnidirectional conversion holds it at zero throughout.
Figure 15: Met3R vs PEGS. We perform DPO finetune ablation on both PEGS and Met3R with our framework and interestingly observe the tail-end distribution improvement offered by PEGS while Met3R barely shows any improvement.