cs.CVSep 29, 2026

Audible World Models: Spatially Aware Sound Generation for 3D Worlds

Authors: Duowen Chen, Jinjin He, Gouthaman KV, Sandeep Bangalore Venkatesh, Bo Zhu

Organizations: Georgia Institute of Technology · Dolby Laboratories

Abstract

Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

    Aug 1, 2026Masaki Yoshida, Ren Togo, Takahiro Ogawa +1Spatial Audio3D Gaussian

  2. Enabling Immersive Audio-Visual Experience from Any Video

    Sep 28, 2026Zitong Lan, Mutian Tong, Jiatao Gu +1Spatial AudioImmersion

  3. Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

    May 27, 2026Jiahao Mei, Heinrich Dinkel, Yadong Niu +7Modern Generative Audio ModelsAcoustic Latent Space