Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.
Figures & tables
Figure 1 : To the best of our knowledge, StereoBind is the first framework for track-conditioned stereo VA generation, with StereoWorld-29K as the first large-scale dataset and StereoWorldBench as the first benchmark for this task.
Figure 2 : Construction pipeline of StereoWorld-29K. We combine curated real-world videos with synthetic audiovisual data. The VAS Pipeline localizes and tracks sound sources and renders track-guided stereo audio, while the SDS Pipeline generates diverse track-controlled scenes.
Figure 3 : Overview of the StereoBind framework. StereoBind introduces VMB for audiovisual source binding, STE for absolute track conditioning, and RT-RoPE for relative motion modeling.
Visual Quality
Audio Quality
Stereo Spatial Fidelity
Model Type
Method
Subject Cons. ↑
Motion Smooth. ↑
Imaging Qual. ↑
Audio PQ ↑
AV Sync ↓
ILD -W ↓
SELD -Acc ↑
SMR -Err ↓
AST -Ang ↑
AST -Cal ↑
VA Models
Ovi [ 22 ]
0.962
0.993
0.681
5.70
0.278
3.744
0.571
65.58
0.196
0.0261
MiniMax-H3
0.968
0.996
0.690
6.81
0.301
3.435
0.725
24.64
0.257
0.131
LTX-2.5 [ 9 ]
0.950
0.995
0.655
6.70
0.367
3.565
0.571
37.62
0.237
0.122
LTX-2.3 [ 9 ]
0.953
0.995
0.653
6.68
0.374
3.574
0.617
39.34
0.221
0.119
V2SA Models
PrismAudio [ 20 ]
–
–
–
6.02
0.479
3.426
0.473
44.55
0.189
0.036
Table 1 : Comparison results on joint video-audio generation. Best results are highlighted in bold, and second-best results are underlined.
Variant
ILD -W ↓
SELD -Acc ↑
SMR -Err ↓
AST -Ang ↑
AST -Cal ↑
w/o Residual Track RoPE
2.800
0.653
19.28
0.308
0.219
w/o Spatial Track Encoder
3.336
0.669
27.38
0.282
0.198
w/o VMB
3.079
0.506
31.16
0.256
0.216
VMB Tokens ( NVMB=16 )
2.785
0.627
18.65
0.322
0.245
VMB Tokens ( NVMB=32 )
2.751
0.691
20.47
0.311
0.246
VMB Tokens ( NVMB=128 )
2.756
0.747
13.47
0.471
0.321
Table 2 : Ablation study of StereoBind on StereoWorldBench (SWBench). Best results are highlighted in bold, and second-best results are underlined.
Figure 4 : Qualitative experimental results on SWBench. We compare StereoBind with open-source VA models and V2SA models. The sound source is highlighted with a red bounding box.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Motion family
Weight
Horizontal left-to-right
0.20
Horizontal right-to-left
0.20
Diagonal left-to-right
0.12
Diagonal right-to-left
0.12
Curved left-to-right
0.12
Curved right-to-left
0.12
Appendix
Table 3: Sampling weights for dynamic motion families.
Motion
Extent
Semantic families
Static
Point-like
stationary animal, small household device, public device, small mechanical device, tonal object
Static
Area-like
large household appliance, water source, large mechanical machine, airflow machine, fire/flame source
Dynamic
Point-like
moving animal, wheeled object, compact vehicle, mobile robot, small mobile machine
Dynamic
Area-like
large vehicle, floor-cleaning machine, outdoor machine, industrial mobile machine
Appendix
Table 4: Semantic-family groups used for different spatial configurations.
Figure 5 : Instructions and output formats of the Qwen-based modules used in StereoWorld-29K construction.
Figure 6 : Dataset statistics of StereoWorld-29K. (a) Distribution of semantic categories of visible sound-producing entities. (b) Composition of real-world and synthetic data sources.
Figure 7 : Motion and spatial statistics of StereoWorld-29K. Dynamic samples are categorized into four motion levels according to horizontal source displacement, while static samples are grouped by left, center, and right source locations.
Method
Immersion Score ↑
Ovi [ 22 ]
1.2
MiniMax-H3
2.1
LTX-2.5 [ 9 ]
2.3
PrismAudio [ 20 ]
1.5
See2Sound [ 5 ]
1.1
StereoBind (Ours)
4.2
Appendix
Table 5 : Human perceptual evaluation on SWBench. Participants rate the overall immersive experience on a five-point Likert scale, considering audiovisual correspondence, spatial motion consistency, stereo realism, and sense of presence. Higher is better.
Figure 8 : Qualitative experimental results on SWBench. We compare StereoBind with open-source VA models and V2SA models. The sound source is highlighted with a red bounding box.
Figure 9 : Qualitative experimental results on SWBench. We compare StereoBind with open-source VA models and V2SA models. The sound source is highlighted with a red bounding box.
Setting
ILD -W ↓
SELD -Acc ↑
SMR -Err ↓
AST -Ang ↑
AST -Cal ↑
Consistent Spatial Conditions
2.754
0.757
14.86
0.465
0.314
Reversed Track Conflict
2.925
0.651
20.78
0.416
0.322
Appendix
Table 6 : Conflicting spatial condition analysis on SWBench. We reverse the numerical motion track while keeping its sparse visual motion reference unchanged, creating contradictory spatial conditions. ↑ indicates higher is better, while ↓ indicates lower is better.
Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.
Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source methods mainly rely on either dual-tower designs with posterior alignment or fully unified tri-modal designs that mix textual context, audio and video in one shared space. The former weakens fine-grained audio-video co-evolution, while the latter couples semantic conditioning with low-level synchronization. To address these limitations, we propose NAVA, a Native Audio-Visual Alignment framework for joint audio-video generation. NAVA is built upon context-conditioned native audio-visual alignment: it first establishes audio-video correspondence in a dedicated interaction space, and then uses external context to condition the joint denoising process. Specifically, NAVA is instantiated with an Align-then-Fuse MMDiT architecture, which transitions from modality-aware audio-video alignment to modality-shared joint denoising. Furthermore, we introduce Timbre-in-Context Conditioning to associate reference timbre cues with corresponding speech spans to achieve controllable speech timbre. Experiments on Verse-Bench and Seed-TTS, together with a user study, demonstrate that NAVA achieves superior video quality, precise audio-visual synchronization, competitive audio quality, and stronger reference-timbre controllability using only 6.3B parameters.