Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Figures & tables
Figure 1: Overview of the proposed Audible World Models pipeline. The numbered blocks correspond to the stages described in Section 3 . Stages 1–6 construct the sound-aware world proxy from a text prompt: a panorama, foreground audible objects, cleaned background regions, segmentation masks, depth, and layered 3D geometry. Stages 7–8 convert semantic sound labels into dry source assets and ground them as persistent 3D sources. Stages 9–11 predict acoustic parameters, propagate each source through the reconstructed geometry with GSound, and output synchronized spatial audio for the moving listener.
Table 2
Method
Mean MOS ↑
Mean rank ↓
Rank-1 ↑
Top-2 ↑
Borda ↑
Stable Audio 1.0
2.98
2.59
21.2%
48.5%
3.41
MMAudio
2.95
2.68
24.2%
47.0%
3.32
See-2-Sound
2.93
3.33
7.6%
24.2%
2.67
OmniAudio
2.52
4.38
1.5%
9.1%
1.62
Ours
3.01
2.02
45.5%
71.2%
3.98
Table 3: Automated VLM-based evaluation using uniformly truncated 5 s clips, including all five methods. MOS values are model-assigned ratings, not human ratings. Higher is better except for mean rank.
Figure 2: Synchronized 20-second audiovisual examples. Compared with scene-level text-to-audio and video-to-audio baselines, our method produces clearer motion-dependent changes because each sound remains tied to a persistent 3D source location.
Figure 3: Synchronized 40-second audiovisual examples. The spatial behavior remains stable over longer trajectories, supporting the claim that explicit source placement and acoustic propagation help maintain long-horizon consistency.
Figure 6
Spatial metrics
Semantic metrics
Variant
ILD ↑
ILD R2↑
Pair Acc. ↑
Spearman ↑
CLAP ↑
IB-Text ↑
IB-Image ↑
Caption ↑
Full model
0.888
0.807
0.823
0.758
0.298
0.191
0.179
0.301
No acoustic simulation
0.007
≪0
0.480
0.035
0.328
0.176
0.177
0.275
No agentic acoustic parameters
0.881
0.753
0.690
0.581
0.287
0.172
0.176
0.204
No semantic decomposition
–
–
–
–
0.324
0.177
0.126
0.246
No geometry-aware placement
0.033
≪0
0.565
0.154
0.290
0.183
0.177
0.267
Table 5: World-sound ablation study. Spatial metrics are computed from localized per-source renders using the original scene geometry and a fixed ILD calibration. Semantic metrics are computed from the final full-scene soundtrack. “–” denotes cases where localized spatial metrics are not applicable.
Method
CLAP ↑
IB-Text ↑
IB-Image ↑
Caption ↑
Weak heuristic
0.121
0.129
0.087
0.129
Strong heuristic
0.225
0.154
0.169
0.270
Full method
0.303
0.196
0.188
0.301
Table 6: Semantic comparison with heuristic world-sound construction. Higher is better. This controlled comparison is reported separately from the original 80-scene benchmark and component-removal ablations.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ILD Corr. ↑
ILD R2↑
Pairwise Rank Acc. ↑
Spearman Score ↑
Ours (20 sec)
0.8994
0.8110
0.8346
0.8063
Ours (40 sec)
0.8875
0.8066
0.8232
0.7577
Appendix
Table 7: Spatial consistency under different trajectory lengths. ILD-based metrics measure agreement between the measured ILD curve and the fitted proxy ILD curve. Energy-ranking metrics compare the predicted source-energy ordering against the distance-based ordering. Higher is better for all metrics.
Variant
CLAP
IB-T
IB-I
Caption
ILD Corr.
ILD R2
Full model
0.303
0.196
0.188
0.301
0.889
0.812
Audio → TangoFlux
0.222
0.167
0.189
0.205
0.796
0.740
Segmentation → ZIM
0.280
0.180
0.204
0.255
0.885
0.791
VLM → Sonnet 4
0.300
0.171
0.222
0.224
0.875
0.776
Depth → DA-V2
0.316
0.185
0.215
0.336
0.854
0.746
VLM + segmentation
0.321
0.180
0.209
0.279
0.876
0.788
Appendix
Table 8: Controlled single- and multi-module replacement. All metrics are higher-is-better. Caption is the Audio Flamingo 2 caption-similarity metric (AF2 in the experiment record). Joint variants retain Stable Audio. These results have their own reported full-model reference and do not replace the original benchmark or component-removal results.
Component
Mean
Median
SD
Min
25th pct.
75th pct.
Max
Foreground labels
1.17
1
0.55
0
1
1.25
2
Background labels
4.20
4
1.53
1
3
5
8
Ambience tracks
1.00
1
0.00
1
1
1
1
Appendix
Table 9: Distribution of semantic audio components per scene. These counts describe labels/tracks, not sampled acoustic emitters.
Parser
Semantic-label agreement
Directional-flag agreement
Foreground
74.2±16.5
82.6±24.5
[67.5,81.5]
[71.4,92.5]
Background
73.3±12.1
88.1±10.0
[68.1,78.3]
[83.7,92.2]
Appendix
Table 10: Parser repeatability over five runs per evaluated scene. Entries are mean ± scene SD; bracketed values are 95% bootstrap confidence intervals. All values are percentages.
Parameter
Mean absolute change
Mean % of allowed span
Reflectivity
0.0217
2.17%
Scattering
0.0688
6.88%
Render volume
0.0210
1.11%
Source volume
3.71
7.41%
Source radius ratio
0.00197
3.95%
Source power
1.13 dB
7.53%
Appendix
Table 11: Acoustic parameter variation between two runs of the same session, averaged over the reported sessions. The final column normalizes the absolute change by the parameter’s allowed span.
Figure 14
Metric
Mean
95% CI
CLAP
0.317
[0.286,0.348]
IB-Text
0.220
[0.130,0.260]
IB-Image
0.262
[0.142,0.302]
Caption
0.276
[0.230,0.322]
Appendix
Table 13: Semantic scores with bootstrap 95% confidence intervals over the original 80 scenes.
Method
ILD Corr.
ILD R2
Mean ILDabs
See-2-Sound
0.1763
0.0408
1.6938
OmniAudio
0.3050
0.1270
2.9048
Stable Audio 1.0
0.1393
0.0291
1.2081
MMAudio
–
–
–
Ours
0.8958
0.8097
9.4735
Appendix
Table 14: Original ILD diagnostics. Correlation and R2 measure agreement with the fitted geometry-based proxy. Mean absolute ILD (dB) measures lateralization strength, not directional accuracy by itself.
Figure 8: Qualitative spatial analysis for one generated scene. We visualize the panorama, the top-down source layout, the proxy ILD curve induced by listener–source geometry, the measured ILD from the rendered binaural audio, and channel-wise spectrograms. The measured ILD follows the geometric trend, indicating that the soundtrack changes consistently with listener motion.
Metric
Estimate
95% CI
MAE ↓
6.91∘
[3.51∘,12.04∘]
Median error ↓
2.96∘
[2.64∘,3.31∘]
Accuracy@ 15∘↑
96.1%
[92.8%,98.9%]
Appendix
Table 15: Full-circle isolated-source DoA recovery for our method. Angular errors are in degrees; accuracy is a percentage.
Figure 9: Runtime distribution of the full generation pipeline for 20s and 40s trajectories.
Operation
20s
40s
Sec.
Share
Sec.
Share
Panorama generation
76.50
8.3%
80.00
8.4%
GPT calls
80.04
8.6%
71.74
7.6%
Ground/segment/inpaint
263.00
28.4%
256.50
27.1%
3D reconstruction/export
323.41
34.9%
334.06
35.2%
Source audio
53.00
5.7%
55.50
5.9%
Appendix
Table 16: Full runtime breakdown of our generation pipeline. The dominant costs are perception and 3D world reconstruction/export, while source-audio generation and acoustic rendering account for a smaller fraction of the total runtime.
Duration (s)
Reported time (s)
5
21.36±9.86
10
34.42±10.61
20
46.34±13.45
40
57.67±19.39
80
106.20±31.77
160
184.14±18.08
Appendix
Table 17: Controlled cached-world scaling. Left: 32 sources with varying trajectory duration. Right: a fixed 20 s trajectory with varying source cap. Times are in seconds; the ± terms reproduce the timing variability reported with the experiments. Source caps concern rendering emitters, not semantic labels.
Moving sources
Warm render (s)
Paired slowdown
Memory (MiB)
1
24.51
1.00×
544.5
2
51.81
2.12×
558.2
4
65.98
2.66×
586.4
8
77.86
3.22×
641.4
16
89.95
3.67×
753.6
32
106.09
4.33×
978.6
Appendix
Table 18: Moving-source rendering with supplied 20 s source and listener trajectories in a static reconstructed world. Warm-render times and paired slowdowns are reported as measured; memory is in MiB. This is a rendering-stage proof of capability, not end-to-end dynamic-world construction.