Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.
Figures & tables
Figure 1 : To the best of our knowledge, StereoBind is the first framework for track-conditioned stereo VA generation, with StereoWorld-29K as the first large-scale dataset and StereoWorldBench as the first benchmark for this task.
Figure 2 : Construction pipeline of StereoWorld-29K. We combine curated real-world videos with synthetic audiovisual data. The VAS Pipeline localizes and tracks sound sources and renders track-guided stereo audio, while the SDS Pipeline generates diverse track-controlled scenes.
Figure 3 : Overview of the StereoBind framework. StereoBind introduces VMB for audiovisual source binding, STE for absolute track conditioning, and RT-RoPE for relative motion modeling.
Visual Quality
Audio Quality
Stereo Spatial Fidelity
Model Type
Method
Subject Cons. ↑
Motion Smooth. ↑
Imaging Qual. ↑
Audio PQ ↑
AV Sync ↓
ILD -W ↓
SELD -Acc ↑
SMR -Err ↓
AST -Ang ↑
AST -Cal ↑
VA Models
Ovi [ 22 ]
0.962
0.993
0.681
5.70
0.278
3.744
0.571
65.58
0.196
0.0261
MiniMax-H3
0.968
0.996
0.690
6.81
0.301
3.435
0.725
24.64
0.257
0.131
LTX-2.5 [ 9 ]
0.950
0.995
0.655
6.70
0.367
3.565
0.571
37.62
0.237
0.122
LTX-2.3 [ 9 ]
0.953
0.995
0.653
6.68
0.374
3.574
0.617
39.34
0.221
0.119
V2SA Models
PrismAudio [ 20 ]
–
–
–
6.02
0.479
3.426
0.473
44.55
0.189
0.036
Table 1 : Comparison results on joint video-audio generation. Best results are highlighted in bold, and second-best results are underlined.
Variant
ILD -W ↓
SELD -Acc ↑
SMR -Err ↓
AST -Ang ↑
AST -Cal ↑
w/o Residual Track RoPE
2.800
0.653
19.28
0.308
0.219
w/o Spatial Track Encoder
3.336
0.669
27.38
0.282
0.198
w/o VMB
3.079
0.506
31.16
0.256
0.216
VMB Tokens ( NVMB=16 )
2.785
0.627
18.65
0.322
0.245
VMB Tokens ( NVMB=32 )
2.751
0.691
20.47
0.311
0.246
VMB Tokens ( NVMB=128 )
2.756
0.747
13.47
0.471
0.321
Table 2 : Ablation study of StereoBind on StereoWorldBench (SWBench). Best results are highlighted in bold, and second-best results are underlined.
Figure 4 : Qualitative experimental results on SWBench. We compare StereoBind with open-source VA models and V2SA models. The sound source is highlighted with a red bounding box.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Motion family
Weight
Horizontal left-to-right
0.20
Horizontal right-to-left
0.20
Diagonal left-to-right
0.12
Diagonal right-to-left
0.12
Curved left-to-right
0.12
Curved right-to-left
0.12
Appendix
Table 3: Sampling weights for dynamic motion families.
Motion
Extent
Semantic families
Static
Point-like
stationary animal, small household device, public device, small mechanical device, tonal object
Static
Area-like
large household appliance, water source, large mechanical machine, airflow machine, fire/flame source
Dynamic
Point-like
moving animal, wheeled object, compact vehicle, mobile robot, small mobile machine
Dynamic
Area-like
large vehicle, floor-cleaning machine, outdoor machine, industrial mobile machine
Appendix
Table 4: Semantic-family groups used for different spatial configurations.
Figure 5 : Instructions and output formats of the Qwen-based modules used in StereoWorld-29K construction.
Figure 6 : Dataset statistics of StereoWorld-29K. (a) Distribution of semantic categories of visible sound-producing entities. (b) Composition of real-world and synthetic data sources.
Figure 7 : Motion and spatial statistics of StereoWorld-29K. Dynamic samples are categorized into four motion levels according to horizontal source displacement, while static samples are grouped by left, center, and right source locations.
Method
Immersion Score ↑
Ovi [ 22 ]
1.2
MiniMax-H3
2.1
LTX-2.5 [ 9 ]
2.3
PrismAudio [ 20 ]
1.5
See2Sound [ 5 ]
1.1
StereoBind (Ours)
4.2
Appendix
Table 5 : Human perceptual evaluation on SWBench. Participants rate the overall immersive experience on a five-point Likert scale, considering audiovisual correspondence, spatial motion consistency, stereo realism, and sense of presence. Higher is better.
Figure 8 : Qualitative experimental results on SWBench. We compare StereoBind with open-source VA models and V2SA models. The sound source is highlighted with a red bounding box.
Figure 9 : Qualitative experimental results on SWBench. We compare StereoBind with open-source VA models and V2SA models. The sound source is highlighted with a red bounding box.
Setting
ILD -W ↓
SELD -Acc ↑
SMR -Err ↓
AST -Ang ↑
AST -Cal ↑
Consistent Spatial Conditions
2.754
0.757
14.86
0.465
0.314
Reversed Track Conflict
2.925
0.651
20.78
0.416
0.322
Appendix
Table 6 : Conflicting spatial condition analysis on SWBench. We reverse the numerical motion track while keeping its sparse visual motion reference unchanged, creating contradictory spatial conditions. ↑ indicates higher is better, while ↓ indicates lower is better.