Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo
Figures & tables
Figure 1: OmniDream transforms a silent perspective video into a 360∘ video with synchronized spatial audio. As users turn their head and inspect different parts of the scene, both the visual perspective and perceived sound change consistently with the environment.
Method
Input
Visual Output
Audio Output
Audio-Visual Synchronization
Physics-Aware Rendering
Multi-Track Audio
MMAudio ( Cheng et al., 2025 )
Video
Video
Mono
✓
✗
✗
Sonic4D ( Xie et al., 2026 )
Video
Video
Stereo
✓
✗
✗
StereoFoley ( Karchkhadze et al., 2025 )
Video
Video
Stereo
✓
✗
✗
ViSAGe ( Kim et al., 2025 )
Video
Video
FOA
✓
✗
✗
OmniAudio ( Liu et al., 2025 )
360∘ video
360∘ video
FOA
✓
✗
✗
See2Sound ( Dagli et al., 2025 )
Image
Image
5.1 Sound
✗
✗
✓
Table 1: Comparison of immersive media generation methods. First-Order Ambisonics (FOA) is a spatial audio format with direction-dependent sound perception. Physics-Aware Rendering indicates whether the method models geometry-based acoustic propagation and environmental effects. Multi-Track Audio indicates whether the method decomposes the scene into individual sounding objects.
Figure 2: OmniDream pipeline . We first expand the narrow perspective video to 360∘ video and reconstruct the scene geometry (Sec. 3.1 ). Based on this, we analyze the sounding object in the scene and track their positions through the video and generate per-object dry sound (Sec. 3.2 ). Finally, we estimate the acoustic properties on the extracted scene mesh and perform ray tracing to simulate the acoustic effect for each object independently and render the final spatial audio (Sec. 3.3 ).
Figure 3: Object tube keyframes and the STFT spectrogram of the generated audio for a swimmer, showing accurate audio-visual timing alignment.
Figure 4: An example of scene geometry and material segmentation. Left is the raw mesh and right is material-annotated mesh, where colors indicate surface regions with different acoustic properties assigned by the VLM.
Figure 6
AV Alignment
Spatial Correctness
Overall
1 fps
5 fps
Method
IB ↑
DS ↓
CC ↑
AUC ↑
CC ↑
AUC ↑
CC ↑
AUC ↑
MMAudio
0.25
0.89
0.11
0.55
0.11
0.52
0.11
0.52
MMAudio (Spatial)
0.24
0.93
0.25
0.61
0.31
0.59
0.20
0.56
See2Sound
0.06
1.08
0.05
0.35
0.09
0.37
0.10
0.38
ViSAGe
0.15
1.00
0.33
0.60
0.35
0.59
0.22
0.61
Table 2: Evaluation of video-to-spatial-audio generation on ImmerseSet.
Figure 7: Spatial audio energy maps overlaid on equirectangular panoramic frames. Warmer colors indicate higher directional energy. OmniDream concentrates energy on visible sounding objects, while baseline models produce incorrect spatial distribution.
AV Alignment
Spatial Correctness
Audio Similarity
Overall
1 fps
5 fps
Method
IB ↑
DS ↓
CC ↑
AUC ↑
CC ↑
AUC ↑
CC ↑
AUC ↑
KL ↓
FD ↓
MMAudio
0.23
0.76
0.17
0.56
0.23
0.57
0.27
0.58
2.09
8.59
MMAudio (Spatial)
0.22
0.79
0.41
0.62
0.44
0.64
0.41
0.61
2.07
8.60
See2Sound
0.05
1.31
0.01
0.44
0.02
0.42
0.02
0.42
2.96
18.53
ViSAGe
0.14
1.21
0.38
0.61
0.36
0.64
0.33
0.63
2.05
11.28
Table 3: Evaluation of video-to-spatial-audio generation on the YT360 dataset.
AV Alignment
Spatial Correctness
Overall
1 fps
5 fps
Variant
IB ↑
DS ↓
CC ↑
AUC ↑
CC ↑
AUC ↑
CC ↑
AUC ↑
w/o object tracking
0.25
0.89
0.26
0.52
0.22
0.47
0.19
0.44
w/o mask dilation
0.23
1.03
0.40
0.71
0.44
0.65
0.27
0.57
Spatialization (replacing acoustic rendering):
azimuth panning
0.21
0.96
0.48
0.71
0.42
0.64
0.37
0.57
Table 4: Ablation studies on the ImmerseSet dataset.
Figure 11
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
PSNR ↑
SSIM ↑
FVD ↓
CubeComposer ( Li et al., 2026 ) (ours)
11.57
0.37
2417
Argus ( Luo et al., 2025 )
11.13
0.34
2974
PanoWan ( Xia et al., 2025 )
8.19
0.22
5099
Appendix
Table 9: Panoramic expansion quality on YT360 clips.
Figure 9: Comparison of perspective-to- 360∘ video generation methods on two scenes. Each row shows the original perspective input (left) alongside the panoramic outputs of CubeComposer ( Li et al., 2026 ) , Argus ( Luo et al., 2025 ) , and PanoWan ( Xia et al., 2025 ) . CubeComposer generates the most coherent panoramic expansion, while Argus and PanoWan exhibit more visible inconsistencies in the generated regions.
Figure 10: Example depth estimation for the generated panoramic videos.
Figure 11: Interactive viewer for the user study. Participants drag to rotate the 360∘ panoramic view; FOA audio is decoded to binaural stereo in real time as the viewpoint changes. The debug overlay (bottom right) shows live decoded stereo RMS/peak levels, and the current yaw/pitch orientation.
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Duowen Chen, Jinjin He, Gouthaman KV +2
Georgia Institute of Technology · Dolby Laboratories
We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and applies a conditioning-dependent low-rank transformation to the generated latent. The adapter is supervised via decoded-audio interaural level and time difference objectives. Trained on MRSAudio, PLACE improves most metrics over ViSAGe on FAIR-Play and achieves an improved SpatialCLAP score over SpatialSonic on the BEWO-1M Single Static test split. Listener evaluations favor PLACE for video-to-audio and out-of-distribution text-to-audio generation, demonstrating flexible multimodal control and improved spatial consistency.
Tiernon Riesenmy, You Zhang, Gautam Bhattacharya +1
The Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GA, USA · Advanced Technology Group, Dolby Laboratories, San Francisco, CA, USA
Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.
Hanmo Chen, Chengcheng Liu, Tianxiao Chen +7
Xidian University · vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.