cs.SDSep 28, 2026

Enabling Immersive Audio-Visual Experience from Any Video

Authors: Zitong Lan, Mutian Tong, Jiatao Gu, Mingmin Zhao

Organizations: University of Pennsylvania

Abstract

Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Audible World Models: Spatially Aware Sound Generation for 3D Worlds

    Sep 29, 2026Duowen Chen, Jinjin He, Gouthaman KV +2Spatial AudioModern Generative Audio Models

  2. PLACE: Positional Latent Adaptation via Conditioned Embeddings for Binaural Audio Generation

    Sep 30, 2026Tiernon Riesenmy, You Zhang, Gautam Bhattacharya +1Audio UnderstandingSound Source Localization

  3. Here the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

    Sep 30, 2026Hanmo Chen, Chengcheng Liu, Tianxiao Chen +7Audio-Visual ConsistencyAudio-Video Generation