Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Figures & tables
Figure 1: Overview of the proposed Audible World Models pipeline. The numbered blocks correspond to the stages described in Section 3 . Stages 1–6 construct the sound-aware world proxy from a text prompt: a panorama, foreground audible objects, cleaned background regions, segmentation masks, depth, and layered 3D geometry. Stages 7–8 convert semantic sound labels into dry source assets and ground them as persistent 3D sources. Stages 9–11 predict acoustic parameters, propagate each source through the reconstructed geometry with GSound, and output synchronized spatial audio for the moving listener.
Table 2
Method
Mean MOS ↑
Mean rank ↓
Rank-1 ↑
Top-2 ↑
Borda ↑
Stable Audio 1.0
2.98
2.59
21.2%
48.5%
3.41
MMAudio
2.95
2.68
24.2%
47.0%
3.32
See-2-Sound
2.93
3.33
7.6%
24.2%
2.67
OmniAudio
2.52
4.38
1.5%
9.1%
1.62
Ours
3.01
2.02
45.5%
71.2%
3.98
Table 3: Automated VLM-based evaluation using uniformly truncated 5 s clips, including all five methods. MOS values are model-assigned ratings, not human ratings. Higher is better except for mean rank.
Figure 2: Synchronized 20-second audiovisual examples. Compared with scene-level text-to-audio and video-to-audio baselines, our method produces clearer motion-dependent changes because each sound remains tied to a persistent 3D source location.
Figure 3: Synchronized 40-second audiovisual examples. The spatial behavior remains stable over longer trajectories, supporting the claim that explicit source placement and acoustic propagation help maintain long-horizon consistency.
Figure 6
Spatial metrics
Semantic metrics
Variant
ILD ↑
ILD R2↑
Pair Acc. ↑
Spearman ↑
CLAP ↑
IB-Text ↑
IB-Image ↑
Caption ↑
Full model
0.888
0.807
0.823
0.758
0.298
0.191
0.179
0.301
No acoustic simulation
0.007
≪0
0.480
0.035
0.328
0.176
0.177
0.275
No agentic acoustic parameters
0.881
0.753
0.690
0.581
0.287
0.172
0.176
0.204
No semantic decomposition
–
–
–
–
0.324
0.177
0.126
0.246
No geometry-aware placement
0.033
≪0
0.565
0.154
0.290
0.183
0.177
0.267
Table 5: World-sound ablation study. Spatial metrics are computed from localized per-source renders using the original scene geometry and a fixed ILD calibration. Semantic metrics are computed from the final full-scene soundtrack. “–” denotes cases where localized spatial metrics are not applicable.
Method
CLAP ↑
IB-Text ↑
IB-Image ↑
Caption ↑
Weak heuristic
0.121
0.129
0.087
0.129
Strong heuristic
0.225
0.154
0.169
0.270
Full method
0.303
0.196
0.188
0.301
Table 6: Semantic comparison with heuristic world-sound construction. Higher is better. This controlled comparison is reported separately from the original 80-scene benchmark and component-removal ablations.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ILD Corr. ↑
ILD R2↑
Pairwise Rank Acc. ↑
Spearman Score ↑
Ours (20 sec)
0.8994
0.8110
0.8346
0.8063
Ours (40 sec)
0.8875
0.8066
0.8232
0.7577
Appendix
Table 7: Spatial consistency under different trajectory lengths. ILD-based metrics measure agreement between the measured ILD curve and the fitted proxy ILD curve. Energy-ranking metrics compare the predicted source-energy ordering against the distance-based ordering. Higher is better for all metrics.
Variant
CLAP
IB-T
IB-I
Caption
ILD Corr.
ILD R2
Full model
0.303
0.196
0.188
0.301
0.889
0.812
Audio → TangoFlux
0.222
0.167
0.189
0.205
0.796
0.740
Segmentation → ZIM
0.280
0.180
0.204
0.255
0.885
0.791
VLM → Sonnet 4
0.300
0.171
0.222
0.224
0.875
0.776
Depth → DA-V2
0.316
0.185
0.215
0.336
0.854
0.746
VLM + segmentation
0.321
0.180
0.209
0.279
0.876
0.788
Appendix
Table 8: Controlled single- and multi-module replacement. All metrics are higher-is-better. Caption is the Audio Flamingo 2 caption-similarity metric (AF2 in the experiment record). Joint variants retain Stable Audio. These results have their own reported full-model reference and do not replace the original benchmark or component-removal results.
Component
Mean
Median
SD
Min
25th pct.
75th pct.
Max
Foreground labels
1.17
1
0.55
0
1
1.25
2
Background labels
4.20
4
1.53
1
3
5
8
Ambience tracks
1.00
1
0.00
1
1
1
1
Appendix
Table 9: Distribution of semantic audio components per scene. These counts describe labels/tracks, not sampled acoustic emitters.
Parser
Semantic-label agreement
Directional-flag agreement
Foreground
74.2±16.5
82.6±24.5
[67.5,81.5]
[71.4,92.5]
Background
73.3±12.1
88.1±10.0
[68.1,78.3]
[83.7,92.2]
Appendix
Table 10: Parser repeatability over five runs per evaluated scene. Entries are mean ± scene SD; bracketed values are 95% bootstrap confidence intervals. All values are percentages.
Parameter
Mean absolute change
Mean % of allowed span
Reflectivity
0.0217
2.17%
Scattering
0.0688
6.88%
Render volume
0.0210
1.11%
Source volume
3.71
7.41%
Source radius ratio
0.00197
3.95%
Source power
1.13 dB
7.53%
Appendix
Table 11: Acoustic parameter variation between two runs of the same session, averaged over the reported sessions. The final column normalizes the absolute change by the parameter’s allowed span.
Figure 14
Metric
Mean
95% CI
CLAP
0.317
[0.286,0.348]
IB-Text
0.220
[0.130,0.260]
IB-Image
0.262
[0.142,0.302]
Caption
0.276
[0.230,0.322]
Appendix
Table 13: Semantic scores with bootstrap 95% confidence intervals over the original 80 scenes.
Method
ILD Corr.
ILD R2
Mean ILDabs
See-2-Sound
0.1763
0.0408
1.6938
OmniAudio
0.3050
0.1270
2.9048
Stable Audio 1.0
0.1393
0.0291
1.2081
MMAudio
–
–
–
Ours
0.8958
0.8097
9.4735
Appendix
Table 14: Original ILD diagnostics. Correlation and R2 measure agreement with the fitted geometry-based proxy. Mean absolute ILD (dB) measures lateralization strength, not directional accuracy by itself.
Figure 8: Qualitative spatial analysis for one generated scene. We visualize the panorama, the top-down source layout, the proxy ILD curve induced by listener–source geometry, the measured ILD from the rendered binaural audio, and channel-wise spectrograms. The measured ILD follows the geometric trend, indicating that the soundtrack changes consistently with listener motion.
Metric
Estimate
95% CI
MAE ↓
6.91∘
[3.51∘,12.04∘]
Median error ↓
2.96∘
[2.64∘,3.31∘]
Accuracy@ 15∘↑
96.1%
[92.8%,98.9%]
Appendix
Table 15: Full-circle isolated-source DoA recovery for our method. Angular errors are in degrees; accuracy is a percentage.
Figure 9: Runtime distribution of the full generation pipeline for 20s and 40s trajectories.
Operation
20s
40s
Sec.
Share
Sec.
Share
Panorama generation
76.50
8.3%
80.00
8.4%
GPT calls
80.04
8.6%
71.74
7.6%
Ground/segment/inpaint
263.00
28.4%
256.50
27.1%
3D reconstruction/export
323.41
34.9%
334.06
35.2%
Source audio
53.00
5.7%
55.50
5.9%
Appendix
Table 16: Full runtime breakdown of our generation pipeline. The dominant costs are perception and 3D world reconstruction/export, while source-audio generation and acoustic rendering account for a smaller fraction of the total runtime.
Duration (s)
Reported time (s)
5
21.36±9.86
10
34.42±10.61
20
46.34±13.45
40
57.67±19.39
80
106.20±31.77
160
184.14±18.08
Appendix
Table 17: Controlled cached-world scaling. Left: 32 sources with varying trajectory duration. Right: a fixed 20 s trajectory with varying source cap. Times are in seconds; the ± terms reproduce the timing variability reported with the experiments. Source caps concern rendering emitters, not semantic labels.
Moving sources
Warm render (s)
Paired slowdown
Memory (MiB)
1
24.51
1.00×
544.5
2
51.81
2.12×
558.2
4
65.98
2.66×
586.4
8
77.86
3.22×
641.4
16
89.95
3.67×
753.6
32
106.09
4.33×
978.6
Appendix
Table 18: Moving-source rendering with supplied 20 s source and listener trajectories in a static reconstructed world. Warm-render times and paired slowdowns are reported as measured; memory is in MiB. This is a rendering-stage proof of capability, not end-to-end dynamic-world construction.
3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition on a single image or viewpoint, their sound is tied to that observation and cannot stay consistent while a listener moves. We introduce the task of generating a spatially consistent soundscape for a given 3DGS world through auditory grounding, identifying which objects in the world should emit sound and anchoring each to a persistent 3D position, and present Scene2Sound, a training-free framework built on this grounding. From the input world alone, our pipeline selects viewpoints that jointly cover the scene, identifies sound-emitting objects with a vision-language model, and associates the multi-view detections into 3D instances through Gaussian set matching, which measures the overlap between the Gaussian sets that render each detection. Each source then receives generated audio that a standard object-based audio engine spatializes in real time at arbitrary listener poses. We further propose two spatial-consistency metrics, one testing whether rendered audio responds consistently to listener motion and one testing whether the claimed sources are supported by views held out from their placement. On a curated set of generated 3DGS worlds and on 3DGS scenes generated from real-world 360-degree captures, Scene2Sound preserves the audio quality of strong per-viewpoint baselines while remaining spatially consistent where per-viewpoint and single-panorama pipelines do not, and a user study confirms the perceptual benefit. Project page: https://masaki-lmd.github.io/scene2sound/.
Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient fine-grained supervision for real-world mixed audio and limited acoustic representations for modeling concurrent audio components. We present Dasheng AudioGen, a unified framework for generating general mixed-audio scenes from text. Dasheng AudioGen introduces structured multi-view captions, which explicitly decouple complex acoustic scenes into complementary description views, thereby enabling fine-grained control over audio layers. Furthermore, we employ a high-dimensional unified semantic-acoustic representation as the shared latent space. It injects semantic priors that facilitate cross-modal training convergence, while its high-dimensional feature space provides sufficient capacity to disentangle and fuse concurrent audio components effectively. With these designs, a simple flow-matching DiT achieves high-quality end-to-end audio scene generation. We also establish a comprehensive evaluation pipeline for audio scene generation. Experiments demonstrate that Dasheng AudioGen achieves performance approaching real-world recordings in mixed-audio categories, while remaining competitive with specialized models in single-type generation tasks. Demos are available at https://nieeim.github.io/Dasheng-AudioGen-Web/.
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
1X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China · 2MiLM Plus, Xiaomi Inc., Beijing, China