cs.SDOct 6, 2026

WorldSonus: Bringing Sound to Worlds

Authors: Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, +4 more

Organizations: The Hong Kong University of Science and Technology · Noiz AI · MetaX · Shanghai Jiao Tong University

Abstract

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Audible World Models: Spatially Aware Sound Generation for 3D Worlds

    Sep 29, 2026Duowen Chen, Jinjin He, Gouthaman KV +2Spatial AudioModern Generative Audio Models

  2. HelixWorld: A Real-time Interactive Audio-Visual World Model

    Sep 29, 2026Lei Ke, Jiahao Pan, Zeyue Tian +13Video World ModelsSpatial Audio

  3. Soundwich: Video Generation with Layered and Controllable Audio

    Sep 30, 2026Zhuo Ning, AmirHossein Naghi Razlighi, Sagi Polaczek +2Audio-Video GenerationVideo Generation