Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Plücker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.
Figures & tables
Figure 1 : MUGEN dataset overview. We introduce a large-scale panoramic video dataset featuring high-quality, long-duration 360° videos totaling 1,300+ hours at the resolution of at least 4K, paired with rich multi-level annotations including camera trajectories, instance masks, depth maps, natural language captions, and structured semantic labels.
Figure 2 : Data curation pipeline for MUGEN. We collect 2,893 hours of panoramic source videos from YouTube, preprocess continuous footage into consecutive one-minute clips for scalable annotation, annotate each clip with semantic and geometric information using Qwen3-VL ( Yang et al., 2025 ) , GPT-4o ( Hurst et al., 2024 ) , and ViPE ( Huang et al., 2025 ) , filter clips by visual quality, overlays, and trajectory validity, and finally sample MUGEN-HQ for model training. The pipeline yields MUGEN (1,318 hours) and MUGEN-HQ (300 hours).
Figure 3 : Statistics of semantic attributes and camera trajectories in MUGEN. Left: semantic attribute distributions. Right: camera trajectory statistics.
Figure 4 : Architecture of Wan360. Wan360 is a camera-controllable interactive panoramic video generation model towards world exploration conditioned on an input panorama, a text prompt, and a target camera trajectory. To adapt the perspective baseline to panoramic geometry, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for cyclic longitude encoding, ERP-aware padding for seam-continuous VAE features, and random roll yaw for yaw-equivalent trajectory augmentation. A panoramic Plücker embedding converts ERP pixels into spherical rays and injects trajectory conditions through the pretrained control adapter.
Model
FVD ↓
SSIM ↑
LPIPS ↓
Consistency ↑
Quality ↑
Dynamic ↑
PSNR ↑
TransErr ↓
(I) Ablation of the three ERP-aware components.
Wan360 (w/o Periodic Longitude RoPE)
491.2
0.435
0.393
0.912
0.459
0.740
14.58
0.192
Wan360 (w/o ERP-Aware Padding)
478.5
0.446
0.378
0.915
0.457
0.710
14.79
0.188
Wan360 (w/o Random Roll Yaw)
486.9
0.439
0.387
0.912
0.456
0.720
14.65
0.215
Wan360 (Ours, full)
476.6
0.449
0.377
0.915
0.463
0.720
14.86
0.184
(II) Disentangled evaluation of model and data.
Table 1: Experimental results on generated video quality and camera control accuracy. (I) Ablation of the three ERP-aware components in Wan360; (II) disentangled evaluation of model and data; (III) comparison with other panoramic video generation methods. ↑ / ↓ indicates whether higher / lower is better.
Figure 5 : Qualitative examples of videos generated by Wan360. Wan360 is a camera-controllable interactive panoramic video generation model towards world exploration. To show the ability of following camera pose, we annotate the trajectory of generated videos.
Figure 6 : Qualitative comparison of panoramic video generation methods. Wan360 produces panoramic videos with higher visual fidelity, better temporal consistency, and more realistic dynamics.
Figure 7 : Qualitative comparison under target camera trajectories. The left column shows the input panorama and target trajectory, while the right columns compare generated sequences from OmniRoam ( Liu et al., 2026 ) and Wan360.
Figure 8 : Qualitative results of Wan360 under different target trajectories for 10s. Given the same input panorama, Wan360 produces distinct video sequences aligned with different camera motions while maintaining consistent scene appearance.
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360° video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that appear plausible yet exhibit inconsistent depth, broken correspondences, and implausible motion across the spherical surface. We address this gap by framing panoramic video generation as a geometry- and dynamics-consistent latent state modeling problem rather than pure visual synthesis. Building on a pre-trained perspective video world model, we introduce two lightweight regularizers: a depth consistency loss against pseudo ground-truth panoramic depth, and a trajectory consistency loss that supervises the 3D world-frame positions of tracked points across time. We further apply spherical-geometry-aware adaptation to the conditioning and positional encoding. We additionally introduce PanoGeo, a unified geometry-aware panoramic video dataset with consistent depth, trajectory, and prompt annotations across diverse real and synthetic sources, used for both training and stratified evaluation. Experiments show that PanoWorld improves geometric consistency over prior panoramic generation methods while maintaining competitive visual realism, establishing that panoramic video generation must be treated as a geometric modeling problem to support the holistic spatial understanding requirements of embodied AI applications. Code is available at https://github.com/ostadabbas/PanoWorld.
We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environment, while interactive roaming requires a complete surrounding world, persistent geometry, controllable camera motion, and temporally coherent high-fidelity observations. MoVerse addresses this problem by separating world construction from observation rendering. It first expands the input into a gravity-aligned 360∘ panorama with topology-aware diffusion, closing the missing field of view before 3D reasoning. It then lifts the panorama into a persistent 3D Gaussian scaffold using panoramic geometry-aware residual prediction, yielding a dense and directly renderable spatial memory. Finally, a Gaussian-conditioned video renderer translates scaffold renderings along user-specified camera trajectories into photorealistic video. To make this renderer practical for interaction, we train a bidirectional diffusion teacher for high-quality conditional rendering and distill it into a causal autoregressive student for bounded-latency streaming. This design combines the controllability and long-range consistency of explicit 3D representations with the perceptual quality of generative video models. MoVerse supports real-time scene roaming at 8FPS on a single NVIDIA RTX4090 GPU, demonstrating a practical path toward single-image world creation with interactive video output.
Yang Zhou, Ziheng Wang, Yuqin Lu +4
South China University of Technology · Orange Team, Moku Lab, HUJING Digital Media & Entertainment Group · Columbia University +1
Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view trajectories, amplifying cross-view inconsistency and temporal drift. We argue that 360° video generation offers a natural solution: panoramic coverage simplifies trajectory design and provides a strong global context for maintaining coherence. We introduce Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion, a controllable 360° video generation framework that synthesizes high-fidelity videos from sparse 360° inputs. The key idea is an explicit 3D Cache, reconstructed from the input, which serves as a geometric scaffold for any user-defined camera path. This allows the diffusion model to focus on photorealistic texture refinement while the 3D Cache enforces global geometric consistency. Experiments show that Pantheon360 achieves superior visual quality and unmatched geometric coherence, enabling reliable and flexible 360° scene generation for downstream simulation and digital-twin applications.
Ting-Hsuan Chen, Ying-Huan Chen, Tao Tu +10
University of Southern California · National Yang Ming Chiao Tung University · Cornell University +1