Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views. We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion. Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions. Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.
Figures & tables
Figure 1 : Language-guided exploration with SPW-Nav. Top left. a shared history generated from a single panorama. Bottom left. a point cloud reconstructed from the generated frames, showing the motion directions of the three candidate routes in the scene. Right. three different instructions given at the same moment continue this history, and each row shows the frames generated after its two instructions. Every frame shows the 360° panorama with its front view inset.
Figure 2 : Overview of the SPW-Nav pipeline. Navigation. The vision-language backbone localizes the referenced target (box) in the latest panorama Iˉk−1 and outputs a camera trajectory (R,t) . Video generation. Memory tokens at long, mid, and short range and the noisy latents of chunk k enter the diffusion transformer, and pose tokens pk,j express t relative to the reference frame ρk . The chunk is denoised from coarse to fine, with guidance v^ at the coarsest scale. Spherical rotation decoupling. The rotation R bypasses the video backbone (dashed line) and is applied to the rotation-free output by spherical resampling. The last frame Iˉk is fed back for the next instruction (solid line).
Method
Δ Rot. ↓
TDE ↓
RE ↓
RPE R ↓
ATE ↓
MA ↑
PSNR ↑
LPIPS ↓
FVD ↓
NWM [ 4 ]
82.74
84.79
83.87
19.18
23.48
0.064
9.90
N/A
1216.1
OmniRoam [ 28 ]
13.00
62.07
13.46
1.603
4.75
0.256
11.47
0.654
694.9
PanoWorld-C. [ 23 ]
13.19
63.36
13.22
1.517
4.25
0.264
11.29
0.645
638.5
PanoWorld-A. [ 23 ]
12.76
63.03
12.99
1.545
4.13
0.257
11.14
0.633
628.8
SPW-Nav
1.45
56.66
3.02
0.780
6.28
0.305
15.69
0.398
351.6
W/o navigation fine-tuning
0.79
100.74
3.84
1.379
10.34
0.028
15.23
0.413
410.7
Table 1 : Language-guided navigation on SPW-NavBench. Each method receives the same instruction through its own interface. The italic row uses the base vision-language model and is not ranked. In all tables, the best is bold and the second best underlined.
Method
Δ Rot. ↓
TDE ↓
RE ↓
RPE R ↓
ATE ↓
MA ↑
PSNR ↑
LPIPS ↓
FVD ↓
NWM [ 4 ]
82.55
86.19
85.64
19.12
22.89
0.081
10.68
N/A
1176.1
CubeComposer [ 24 ]
13.97
91.40
13.97
1.521
22.60
0.073
14.61
0.565
973.1
OmniRoam [ 28 ]
12.30
77.12
15.07
2.042
10.40
0.111
15.36
0.489
610.5
PanoWorld-C. [ 23 ]
13.46
26.28
13.47
1.528
4.24
0.688
16.12
0.460
462.0
PanoWorld-A. [ 23 ]
13.61
24.26
13.62
1.530
3.93
0.726
16.24
0.425
384.2
Matrix-Game 3.0 + CubeComposer
13.13
92.91
13.14
1.479
20.24
0.080
14.59
0.587
892.8
Table 2 : Camera-controlled panoramic generation from ground-truth trajectories.
Figure 3 : Qualitative comparison with ground-truth trajectories. Each row shows the panoramic video generated by one method at six evenly spaced timestamps from left to right. The last row shows the ground truth, and SPW-Nav is outlined in green.
Method
Memory gain ↑
Closure rot. ( ∘ ) ↓
Closure trans. (%) ↓
OmniRoam [ 28 ]
0.026
11.02
57.5
PanoWorld-A. [ 23 ]
0.084
39.52
59.0
HunyuanWorld-Voyager † [ 19 ]
0.087
8.91
40.7
Matrix-Game 3.0 † [ 43 ]
0.249
28.41
32.7
+ CubeComposer [ 24 ]
0.204
22.26
32.0
SPW-Nav
0.173
1.62
30.7
Table 3 : Long-horizon consistency on 40 closed routes of twelve chunks. † Native perspective view.
Figure 4 : Streaming generation with instruction switching. The first column shows the input panorama, and each instruction above spans the frames generated while following it.
NTL
OTG
TDE ↓
MA ↑
ATE ↓
RE ↓
RPE R ↓
PSNR ↑
LPIPS ↓
FVD ↓
29.15
0.469
4.80
26.58
3.157
13.69
0.494
512.2
✓
27.85
0.517
5.64
28.13
3.350
13.26
0.514
562.3
✓
43.53
0.395
12.33
0.99
0.160
17.05
0.361
395.6
✓
✓
24.30
0.631
5.00
1.84
0.282
16.48
0.382
308.0
W/o SRD
27.99
0.521
5.99
31.31
3.801
12.97
0.518
596.5
Table 4 : Ablation on SPW-NavBench. The upper block varies NTL and OTG with SRD enabled. The lower row disables both SRD and NTL.
Figure 71 : Zero-shot streams generated by SPW-Nav from panoramas outside its training data. Each row is one stream of 14 to 30 seconds with evenly spaced frames.
Figure 72 : One-minute streams from SPW-Nav. Each row spans 60 to 74 seconds.
Test scene
SPW-NavBench
Method
Action
Goal
Action
Goal
Template rules
0.100
0.000
0.375
0.000
W/o fine-tuning
0.663
0.154
0.850
0.450
W/o image
0.900
0.135
1.000
0.000
Ours
0.925
0.692
1.000
0.617
Table 71 : Navigation accuracy. Accuracy of the vision-language backbone on the held-out test scene and on the navigation instructions of SPW-NavBench.
Segment
1
2
3
4
TDE ( ∘ ) ↓
11.2
N/A
31.8
19.1
RE ( ∘ ) ↓
0.9
2.8
1.1
0.6
MA ↑
0.86
1.00
0.51
0.69
FVD ↓
224.9
364.4
316.3
435.5
Table 72 : Streaming generation. Metrics over four consecutive instructions in 20 scenes. Since segment 2 is a pure turn, its TDE is undefined.
Method
PSNR ↑
LPIPS ↓
FVD ↓
TDE ↓
RE ↓
MA ↑
RPE R ↓
Stable Virtual Camera [ 55 ]
12.52
0.531
553.6
26.3
6.95
0.677
1.987
ReCamMaster [ 2 ]
15.54
0.440
353.7
48.9
5.19
0.336
0.836
CamI2V [ 54 ]
16.54
0.364
332.8
30.4
3.29
0.666
0.483
HunyuanWorld-Voyager [ 19 ]
16.57
0.388
301.0
17.1
3.45
0.768
0.583
CameraCtrl [ 16 ]
17.31
0.340
269.8
25.5
3.28
0.713
0.493
Matrix-Game 3.0 [ 43 ]
16.10
0.355
844.2
93.2
8.69
0.057
0.943
Table 73 : Comparison with perspective generators on SPW-NavBench.
NTL
κ
TDE ↓
MA ↑
ATE ↓
RE ↓
RPE R ↓
PSNR ↑
LPIPS ↓
FVD ↓
1
29.15
0.469
4.80
26.58
3.157
13.69
0.494
512.2
2
27.85
0.517
5.64
28.13
3.350
13.26
0.514
562.3
✓
1
43.53
0.395
12.33
0.99
0.160
17.05
0.361
395.6
✓
2
24.30
0.631
5.00
1.84
0.282
16.48
0.382
308.0
✓
3
24.09
0.649
5.25
3.14
0.506
15.67
0.412
333.4
✓
4
24.80
0.642
6.24
5.71
0.959
14.99
0.442
381.5
Table 74 : Full sweep of the OTG weight κ with and without NTL on SPW-NavBench.
Figure 101 : Sample clips of SPW-NavSet. Each row is one clip, with frames evenly spaced from the first to the last frame. The label on the left gives the source of the clips.
Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Plücker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.
Jiaming Tan, Zhen Li, Shuwei Shi +6
Alaya Lab · Beijing Institute of Technology · Shanghai Innovation Institute
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360° video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that appear plausible yet exhibit inconsistent depth, broken correspondences, and implausible motion across the spherical surface. We address this gap by framing panoramic video generation as a geometry- and dynamics-consistent latent state modeling problem rather than pure visual synthesis. Building on a pre-trained perspective video world model, we introduce two lightweight regularizers: a depth consistency loss against pseudo ground-truth panoramic depth, and a trajectory consistency loss that supervises the 3D world-frame positions of tracked points across time. We further apply spherical-geometry-aware adaptation to the conditioning and positional encoding. We additionally introduce PanoGeo, a unified geometry-aware panoramic video dataset with consistent depth, trajectory, and prompt annotations across diverse real and synthetic sources, used for both training and stratified evaluation. Experiments show that PanoWorld improves geometric consistency over prior panoramic generation methods while maintaining competitive visual realism, establishing that panoramic video generation must be treated as a geometric modeling problem to support the holistic spatial understanding requirements of embodied AI applications. Code is available at https://github.com/ostadabbas/PanoWorld.
We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single 360∘ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory control, which hinders reliable downstream 3D reconstruction, or struggle with large disocclusions under long-range camera motion while requiring high-end multi-GPU servers.We present Genie Sim PanoWorld, a two-stage feed-forward pipeline that bridges generation and reconstruction via an explicit, trajectory-controllable panoramic video. A NavMesh-planned SE(3) roaming trajectory is injected into a latent video diffusion model through dense geometry-warped conditioning; long--short trajectory mixed training and a self-consistency objective based on shortcut models together yield high-fidelity video in four CFG-free denoising steps. A feed-forward panoramic reconstructor then lifts the generated video into a high-fidelity 3D Gaussian scene that supports real-time, free-viewpoint roaming and can be directly used as a simulation-ready asset for embodied AI applications. Experiments show that Genie Sim PanoWorld outperforms geometry-conditioned baselines in both panoramic video generation and downstream 3D reconstruction, while generalizing zero-shot to unseen indoor scenes.