Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/
Figures & tables
Figure 1: Overview of WorldSonus. (a) Causal streaming pipeline : Video frames map to two-timescale visual tokens conditioning an AR transformer and a rectified-flow head. (b) Training-only ShiftNCE : A training-only frozen Synchformer teacher guides temporal alignment via contrastive window matching. (c) Spatial stereo supervision : In-the-wild stereo filtering and panoramic FOA view-decoding provide directional audio grounding. (d) Interactive prompt control : Text prompts update dynamically at chunk boundaries, preserving session state and acoustic continuity.
Distribution (mid)
Spatial
Semantic
Temporal
Set
Model
Acc.
Cond.
Chunk/ t
FAD ↓
FD P ↓
FD O ↓
KL P ↓
S-FD O ↓
IB ↑
CLAP ↑
DeSync ↓
VGG 5 s
AudioX
Bi
VT
–
3.00
153.29
38.19
1.66
87.60
26.96
37.87
0.986
ThinkSound
Bi
VT
–
2.94
141.40
47.32
1.91
62.95
24.12
33.62
0.481
PrismAudio
Bi
VT
–
2.09
132.58
50.72
1.67
94.77
26.28
39.20
0.539
V-AURA
Str
V
640 / 636
4.06
304.06
51.01
2.03
–
26.48
24.91
1.287
WorldSonus (Ours)
Caus
VT
100 / 41.2
1.73
159.74
39.03
1.48
43.47
27.06
38.35
0.686
Table 1: Quantitative comparison across five held-out splits ( 4,096 clips for 5 s/ 10 s; 1,024 for 30 s). Best results are in bold, and second-best are underlined. Acc.: Bi (Bidirectional), Str (Stream), Caus (Causal). Cond.: VT (video+text), V (video-only). Chunk/ t : chunk length and compute time in ms. IB and CLAP are multiplied by 100 . S-FD O is omitted for monophonic V-AURA.
Table 3
Figure 2: Qualitative stereo layout versus bidirectional baselines on two interactive clips. Red boxes mark sources. Energy-balance curves show left (negative) versus right (positive) channel dominance.
Protocol
FAD ↓
FD P ↓
FD O ↓
KL P ↓
S-FD O ↓
IB ↑
DeSync ↓
Direct last 5 s
2.63
271.68
35.16
1.22
39.24
21.38
0.880
Rollout tail ( 25 – 30 s)
2.51
246.69
40.46
1.31
41.77
19.99
0.827
Table 4: Long-horizon stability on Interactive 30 s ( 1,024 clips). Final 5 s window ( 25 – 30 s) generated directly versus via continuous rollout.
Model
VGG 10 s
Inter. 10 s
AudioX
15.41
14.04
ThinkSound
13.92
13.43
PrismAudio
16.16
11.99
V-AURA (video-only)
12.16
12.11
WorldSonus (Ours)
25.68
23.44
Table 5: Dual relative match rate (%). Both halves prefer their own reference prompt.
Figure 3: User-study preference matrices across (a) spatial alignment, (b) temporal alignment, (c) semantic alignment, and (d) overall preference. Cell (i,j) is the percentage preferring the row method over the column method, counting each tie as half a vote ( 40 ratings per pair).
Table 8
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Configuration and Architectural Specification
Streaming Unit
100 ms chunk: 3 latent frames and 3 video frames at 30 FPS
Causal Stereo VAE
Frozen SoundReactor ( Saito et al., 2025 ) VAE; maps 48 kHz stereo audio to 30 Hz latents
Frozen DINOv3 S+ ( Siméoni et al., 2026 ) ; aspect-preserving resolution buckets; patch grid pooled 2×
Visual Adapter
Linear bottleneck to 128 channels, concatenated with ΔGn ; 2D-RoPE ( Su et al., 2024 ; Heo et al., 2024 ) ; 1 learned query per frame plus 1 chunk-summary query
Appendix
Table 9: System implementation and architectural configuration of WorldSonus.
Stage
Corpus Composition
Task Ratio (video-audio : audio-only)
View Representation
Pretraining ( 0 – 140 k)
Full video-audio/audio-only mixture: base stereo datasets plus mono-dominant expansion
2:1
5 s blocks, with 10 s paired views on 4 of 5 epochs
Fine-tuning ( 140 – 150 k)
High-synchrony filtered stereo video-audio only (expansion excluded)
1:0
5 s blocks, with 10 s paired views on 4 of 5 epochs
Appendix
Table 10: Training curriculum: corpora, task mixtures, and view lengths.
Linear warmup from 0 to 5×10−5 over 4,000 steps; constant 5×10−5 thereafter
Gradient Clipping
Maximum gradient norm 1.0
EMA Rate
Decay rate 0.9999
Distributed Setup
16 NVIDIA H100 GPUs ( 2 nodes); batch size 16 per GPU ( 256 global batch)
Appendix
Table 11: Training configuration and hyperparameter settings.
Data Pool
5-s Clips
10-s Paired Views
Total Entries
Audio Duration (h)
Base stereo video-audio
590,510
139,079
590,510
820.15
Mono-dominant video-audio
128,885
48,105
128,885
179.01
Base audio-only
66,859
60,832
127,691
261.84
Recovered audio-only
146,834
–
146,834
203.94
Total Training Inventory
933,088
60,832
993,920
1,464.93
Appendix
Table 12: Training corpus breakdown by pool and audio duration.
Set
Clips
Cand. Win.
Elig. Clips
Elig. Win. (L / R)
Keep (%)
VGGSound 10 s
4,096
40,960
1,589
3,314 / 3,356
16.3
Interactive 10 s
4,096
40,960
2,275
6,759 / 3,316
24.6
Interactive 30 s
1,024
30,720
679
2,870 / 3,200
19.8
Total
9,216
112,640
–
12,943 / 9,872
20.3
Appendix
Table 13: Eligible 1 s window counts and retention rates for BiasSkill evaluation.
Set
Model
L Rec.
R Rec.
Bal. Acc.
RMS Cov.
Act. Cov.
Null μnull
Δ
BiasSkill (%)
p
VGG 10 s
AudioX
0.273
0.186
0.229
0.990
0.430
0.199
+0.030
3.79
0.0002
ThinkSound
0.181
0.386
0.283
0.956
0.554
0.260
+0.024
3.22
0.0035
PrismAudio
0.277
0.093
0.185
0.946
0.364
0.176
+0.009
1.10
0.1052
WorldSonus (Ours)
0.256
0.211
0.233
0.976
0.439
0.198
+0.035
4.42
0.0001
Int. 10 s
AudioX
0.256
0.237
0.246
0.967
0.460
0.222
+0.024
3.08
0.0001
ThinkSound
0.159
0.395
0.277
0.983
0.594
0.290
−0.013
−1.78
0.9568
Appendix
Table 14: Detailed stereo balance agreement results under the −60 dBFS generation-RMS gate. L/R Rec. denotes per-side recall; Bal. Acc. is Abal ; RMS Cov. is the coverage of RMSgen≥0.001 ; Act. Cov. requires both RMS and ∣dgen∣≥0.1 ; Δ=Abal−μnull ; BiasSkill represents the normalized effect size (%); and p indicates the one-sided permutation test value.
Set
Arm
s11
s12
s21
s22
S
M
R (%)
VGG 10 s
Switch
0.3823
0.3312
0.3293
0.3530
0.3677
0.0374
25.68
Hold
0.3823
0.3312
0.3543
0.3270
0.3546
0.0119
14.43
Int. 10 s
Switch
0.3560
0.3062
0.3062
0.3284
0.3422
0.0360
23.44
Hold
0.3560
0.3062
0.3316
0.3013
0.3286
0.0097
12.82
Appendix
Table 15: Full paired-control intervention results ( 4,096 clips per set). The sij entries are raw CLAP cosine similarities. S is the mean of the aligned similarities s11 and s22 , M is the mean matching advantage, and R is reported as a percentage (%). Hold is evaluated against the same reference prompt pair (P1,P2) as Switch.
Evaluation Set
Quartile
Text Cosine Range
Intervention Gain G
VGGSound 10 s
Q1 (Most Distinct)
[−0.244,0.525]
0.1393[0.1287,0.1499]
Q2
[0.525,0.727]
0.0441[0.0388,0.0494]
Q3
[0.727,0.842]
0.0157[0.0127,0.0187]
Q4 (Most Similar)
[0.842,0.990]
0.0047[0.0030,0.0064]
Interactive 10 s
Q1 (Most Distinct)
[−0.287,0.490]
0.1277[0.1187,0.1367]
Q2
[0.491,0.688]
0.0514[0.0459,0.0569]
Appendix
Table 16: Stratification of paired prompt gain G across text-similarity quartiles ( 1,024 clips per quartile). Brackets denote 95% clip-level bootstrap percentile intervals.
Figure 4: User-study interface for gameplay and real-world clips. Participants compare samples A and B and select A, Tie, or B for spatial alignment, temporal synchronization, semantic alignment, and overall preference.
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Duowen Chen, Jinjin He, Gouthaman KV +2
Georgia Institute of Technology · Dolby Laboratories
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.
Lei Ke, Jiahao Pan, Zeyue Tian +13
The Hong Kong University of Science and Technology · Noiz AI
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.