cs.CVSep 30, 2026

Here the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Authors: Hanmo Chen, Chengcheng Liu, Tianxiao Chen, Zheyu Zhang, Siming Zheng, Jinwei Chen, Xu Yang, Cheng Deng, +2 more

Organizations: Xidian University · vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.

Abstract

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Enabling Immersive Audio-Visual Experience from Any Video

    Sep 28, 2026Zitong Lan, Mutian Tong, Jiatao Gu +1Spatial AudioImmersion

  2. Soundwich: Video Generation with Layered and Controllable Audio

    Sep 30, 2026Zhuo Ning, AmirHossein Naghi Razlighi, Sagi Polaczek +2Audio-Video GenerationVideo Generation

  3. Native Audio-Visual Alignment for Generation

    May 28, 2026Longbin Ji, Guan Wang, Xuan Wei +6Audio-Video GenerationAudio-Visual Consistency