Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/
Figures & tables
Figure 1: Overview of WorldSonus. (a) Causal streaming pipeline : Video frames map to two-timescale visual tokens conditioning an AR transformer and a rectified-flow head. (b) Training-only ShiftNCE : A training-only frozen Synchformer teacher guides temporal alignment via contrastive window matching. (c) Spatial stereo supervision : In-the-wild stereo filtering and panoramic FOA view-decoding provide directional audio grounding. (d) Interactive prompt control : Text prompts update dynamically at chunk boundaries, preserving session state and acoustic continuity.
Distribution (mid)
Spatial
Semantic
Temporal
Set
Model
Acc.
Cond.
Chunk/ t
FAD ↓
FD P ↓
FD O ↓
KL P ↓
S-FD O ↓
IB ↑
CLAP ↑
DeSync ↓
VGG 5 s
AudioX
Bi
VT
–
3.00
153.29
38.19
1.66
87.60
26.96
37.87
0.986
ThinkSound
Bi
VT
–
2.94
141.40
47.32
1.91
62.95
24.12
33.62
0.481
PrismAudio
Bi
VT
–
2.09
132.58
50.72
1.67
94.77
26.28
39.20
0.539
V-AURA
Str
V
640 / 636
4.06
304.06
51.01
2.03
–
26.48
24.91
1.287
WorldSonus (Ours)
Caus
VT
100 / 41.2
1.73
159.74
39.03
1.48
43.47
27.06
38.35
0.686
Table 1: Quantitative comparison across five held-out splits ( 4,096 clips for 5 s/ 10 s; 1,024 for 30 s). Best results are in bold, and second-best are underlined. Acc.: Bi (Bidirectional), Str (Stream), Caus (Causal). Cond.: VT (video+text), V (video-only). Chunk/ t : chunk length and compute time in ms. IB and CLAP are multiplied by 100 . S-FD O is omitted for monophonic V-AURA.
Table 3
Figure 2: Qualitative stereo layout versus bidirectional baselines on two interactive clips. Red boxes mark sources. Energy-balance curves show left (negative) versus right (positive) channel dominance.
Protocol
FAD ↓
FD P ↓
FD O ↓
KL P ↓
S-FD O ↓
IB ↑
DeSync ↓
Direct last 5 s
2.63
271.68
35.16
1.22
39.24
21.38
0.880
Rollout tail ( 25 – 30 s)
2.51
246.69
40.46
1.31
41.77
19.99
0.827
Table 4: Long-horizon stability on Interactive 30 s ( 1,024 clips). Final 5 s window ( 25 – 30 s) generated directly versus via continuous rollout.
Model
VGG 10 s
Inter. 10 s
AudioX
15.41
14.04
ThinkSound
13.92
13.43
PrismAudio
16.16
11.99
V-AURA (video-only)
12.16
12.11
WorldSonus (Ours)
25.68
23.44
Table 5: Dual relative match rate (%). Both halves prefer their own reference prompt.
Figure 3: User-study preference matrices across (a) spatial alignment, (b) temporal alignment, (c) semantic alignment, and (d) overall preference. Cell (i,j) is the percentage preferring the row method over the column method, counting each tie as half a vote ( 40 ratings per pair).
Table 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Configuration and Architectural Specification
Streaming Unit
100 ms chunk: 3 latent frames and 3 video frames at 30 FPS
Causal Stereo VAE
Frozen SoundReactor ( Saito et al., 2025 ) VAE; maps 48 kHz stereo audio to 30 Hz latents
Frozen DINOv3 S+ ( Siméoni et al., 2026 ) ; aspect-preserving resolution buckets; patch grid pooled 2×
Visual Adapter
Linear bottleneck to 128 channels, concatenated with ΔGn ; 2D-RoPE ( Su et al., 2024 ; Heo et al., 2024 ) ; 1 learned query per frame plus 1 chunk-summary query
Appendix
Table 9: System implementation and architectural configuration of WorldSonus.
Stage
Corpus Composition
Task Ratio (video-audio : audio-only)
View Representation
Pretraining ( 0 – 140 k)
Full video-audio/audio-only mixture: base stereo datasets plus mono-dominant expansion
2:1
5 s blocks, with 10 s paired views on 4 of 5 epochs
Fine-tuning ( 140 – 150 k)
High-synchrony filtered stereo video-audio only (expansion excluded)
1:0
5 s blocks, with 10 s paired views on 4 of 5 epochs
Appendix
Table 10: Training curriculum: corpora, task mixtures, and view lengths.
Linear warmup from 0 to 5×10−5 over 4,000 steps; constant 5×10−5 thereafter
Gradient Clipping
Maximum gradient norm 1.0
EMA Rate
Decay rate 0.9999
Distributed Setup
16 NVIDIA H100 GPUs ( 2 nodes); batch size 16 per GPU ( 256 global batch)
Appendix
Table 11: Training configuration and hyperparameter settings.
Data Pool
5-s Clips
10-s Paired Views
Total Entries
Audio Duration (h)
Base stereo video-audio
590,510
139,079
590,510
820.15
Mono-dominant video-audio
128,885
48,105
128,885
179.01
Base audio-only
66,859
60,832
127,691
261.84
Recovered audio-only
146,834
–
146,834
203.94
Total Training Inventory
933,088
60,832
993,920
1,464.93
Appendix
Table 12: Training corpus breakdown by pool and audio duration.
Set
Clips
Cand. Win.
Elig. Clips
Elig. Win. (L / R)
Keep (%)
VGGSound 10 s
4,096
40,960
1,589
3,314 / 3,356
16.3
Interactive 10 s
4,096
40,960
2,275
6,759 / 3,316
24.6
Interactive 30 s
1,024
30,720
679
2,870 / 3,200
19.8
Total
9,216
112,640
–
12,943 / 9,872
20.3
Appendix
Table 13: Eligible 1 s window counts and retention rates for BiasSkill evaluation.
Set
Model
L Rec.
R Rec.
Bal. Acc.
RMS Cov.
Act. Cov.
Null μnull
Δ
BiasSkill (%)
p
VGG 10 s
AudioX
0.273
0.186
0.229
0.990
0.430
0.199
+0.030
3.79
0.0002
ThinkSound
0.181
0.386
0.283
0.956
0.554
0.260
+0.024
3.22
0.0035
PrismAudio
0.277
0.093
0.185
0.946
0.364
0.176
+0.009
1.10
0.1052
WorldSonus (Ours)
0.256
0.211
0.233
0.976
0.439
0.198
+0.035
4.42
0.0001
Int. 10 s
AudioX
0.256
0.237
0.246
0.967
0.460
0.222
+0.024
3.08
0.0001
ThinkSound
0.159
0.395
0.277
0.983
0.594
0.290
−0.013
−1.78
0.9568
Appendix
Table 14: Detailed stereo balance agreement results under the −60 dBFS generation-RMS gate. L/R Rec. denotes per-side recall; Bal. Acc. is Abal ; RMS Cov. is the coverage of RMSgen≥0.001 ; Act. Cov. requires both RMS and ∣dgen∣≥0.1 ; Δ=Abal−μnull ; BiasSkill represents the normalized effect size (%); and p indicates the one-sided permutation test value.
Set
Arm
s11
s12
s21
s22
S
M
R (%)
VGG 10 s
Switch
0.3823
0.3312
0.3293
0.3530
0.3677
0.0374
25.68
Hold
0.3823
0.3312
0.3543
0.3270
0.3546
0.0119
14.43
Int. 10 s
Switch
0.3560
0.3062
0.3062
0.3284
0.3422
0.0360
23.44
Hold
0.3560
0.3062
0.3316
0.3013
0.3286
0.0097
12.82
Appendix
Table 15: Full paired-control intervention results ( 4,096 clips per set). The sij entries are raw CLAP cosine similarities. S is the mean of the aligned similarities s11 and s22 , M is the mean matching advantage, and R is reported as a percentage (%). Hold is evaluated against the same reference prompt pair (P1,P2) as Switch.
Evaluation Set
Quartile
Text Cosine Range
Intervention Gain G
VGGSound 10 s
Q1 (Most Distinct)
[−0.244,0.525]
0.1393[0.1287,0.1499]
Q2
[0.525,0.727]
0.0441[0.0388,0.0494]
Q3
[0.727,0.842]
0.0157[0.0127,0.0187]
Q4 (Most Similar)
[0.842,0.990]
0.0047[0.0030,0.0064]
Interactive 10 s
Q1 (Most Distinct)
[−0.287,0.490]
0.1277[0.1187,0.1367]
Q2
[0.491,0.688]
0.0514[0.0459,0.0569]
Appendix
Table 16: Stratification of paired prompt gain G across text-similarity quartiles ( 1,024 clips per quartile). Brackets denote 95% clip-level bootstrap percentile intervals.
Figure 4: User-study interface for gameplay and real-world clips. Participants compare samples A and B and select A, Tie, or B for spatial alignment, temporal synchronization, semantic alignment, and overall preference.