World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.
Figures & tables
Figure 1 : HelixWorld. Real-time interaction across diverse worlds with synchronized video and camera-aligned spatial audio.
Source
Hours
Share
Real-world
1.8k
43.9%
Game
1.4k
34.1%
Open-source
0.9k
22.0%
Total
4.1k
100%
Table 1 : Corpus by source.
Figure 2 : Data pipeline. Raw video and audio pass four cost-ordered stages; surviving clips fork into parallel camera recovery and decoupled captioning branches to yield the trainable corpus.
Stage
Hours
Pass Rate
Raw footage
4.1k
100%
Heuristic filtering
3.4k
81.9%
Semantic filtering
3.2k
95.5%
Camera annotation
3.0k
93.5%
Trainable corpus
3.0k
73.1%
Table 2 : Curation yield.
Figure 3 : Progressive condition injection. Three-stage training: optimize control modules with the backbone frozen, unfreeze the video branch, and jointly train the full audio-visual backbone.
Figure 4 : Causal distillation pipeline. Top: Self-forcing updates stochastically alternate between online trajectory distillation and distribution matching. Bottom: Long-horizon tuning extends rollouts segment by segment via a sliding KV cache with initial sink frames and detached history.
Method
Audio
RTF ↓
Average ↑
Quality ↑
Setting ↑
Interaction ↑
Consistency ↑
Physical ↑
Alaya-EVOKE-Turbo
✗
4.29
82.0
81.9
82.1
83.9
88.1
74.0
EchoWM
✓
1.98
81.0
81.1
77.5
87.9
88.3
70.1
Zing-0.5
✗
2.81
81.0
80.6
77.8
84.2
88.5
73.8
LingBot-World v2
✗
5.21
79.4
81.8
76.8
82.8
86.5
69.1
HY-World 1.5
✗
12.33
78.1
78.1
72.2
86.8
86.9
66.3
Lyra 2.0
✗
30.87
76.4
77.1
73.2
85.6
79.3
66.7
Table 3 : Visual quality and action following on the WBench navigation split [ 51 ] . Audio indicates native audio synthesis; RTF is measured on a single NVIDIA H800.
Model
KL ↓
FAD ↓
DeSync (s) ↓
IB ↑
CLAP ↑
Spatial ↑
LTX-2.3 (base)
1.8040
7.6354
0.4042
0.2777
0.2577
0.9826
EchoWM
1.9530
7.7472
0.6745
0.2036
0.2352
12.6476
HelixWorld + AudioX
2.0190
3.9166
1.1764
0.2569
0.3094
−5.7392
HelixWorld + ThinkSound
2.2335
6.1876
0.4236
0.2079
0.2888
14.8467
HelixWorld + PrismAudio
2.2427
6.3873
0.4582
0.2152
0.2716
−9.5072
HelixWorld (bidirectional teacher)
1.5414
3.0775
0.4988
0.2700
0.2676
33.2428
Table 4 : Audio-visual quality on HelixBench . HelixWorld + X re-dubs our video using V2A model X; joint generation achieves superior spatial acoustic alignment.
Table 9
Ltraj
Long-Horizon
Average ↑
Quality ↑
Setting ↑
Interaction ↑
Consistency ↑
Physical ↑
✗
✗
77.9
77.0
75.3
85.0
85.3
66.9
✓
✗
78.7
77.8
74.1
86.1
86.6
68.9
✗
✓
78.4
79.0
77.5
86.5
83.7
65.1
✓
✓
79.1
79.3
77.1
86.8
84.6
67.7
Table 7 : Ablation on self-forcing and long-horizon tuning. Trajectory distillation ( Ltraj ) and long-horizon tuning offer complementary gains, jointly reaching the best overall WBench score.
Figure 6 : Visual comparison under interactive control. HelixWorld retains scene structure during navigation (left) and motion reversal (right). Key overlays indicate controls.
Figure 7 : User-study preferences. Purple indicates preference for HelixWorld, gray a tie, and beige preference for the baseline across full audio-visual rollouts (a) and muted video dynamics (b). Dashed vertical lines mark 50%.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Generator / fake-score adaptation
LoRA rank 256 , alpha 256 , dropout 0
Optimizer
AdamW, β1=0.9 , β2=0.999
Generator / fake learning rate
10−5 , constant; no warmup or decay
Gradient norm clipping
1.0
Precision
BF16 with gradient checkpointing
Global batch size
48 ; one sample per GPU, no accumulation
Appendix
Table 9 : Self-forcing hyperparameters. Settings shared by DMD-only self-forcing and self-forcing with online trajectory distillation loss.
Source Family
Pose Origin
Raw Hours
Share
Trainable Clips
Real-world web video
Estimated (VGGT- Ω + DA3)
1.8 k
43.9%
∼0.91 M
Game screen recordings
Engine-logged (ground truth)
1.4 k
34.1%
∼0.73 M
Open-source subsets
Mixed (logged / estimated)
0.9 k
22.0%
∼0.46 M
Total Raw Footage
—
4.1 k
100%
∼2.93 M
Trainable Corpus
Verified metric poses
3.0 k
73.1%
2.10 M
Appendix
Table 10 : Data sources, domain distribution, and corpus statistics. Pose origin indicates whether camera trajectories are logged directly from engine states or reconstructed offline via geometric estimation. All statistics reflect the frozen corpus.
Method
Translation (%) ↓
Rotation ( ∘ ) ↓
Runtime (s) ↓
L1
RMSE
VGGT [ 41 ]
1.010
0.760
0.315
14.23
VGGT-Omega [ 42 ]
1.001
0.768
0.234
16.09
ViPE [ 17 ]
0.763
0.582
0.263
187.32
Appendix
Table 11 : Mean camera estimation errors and runtime across benchmark scenes. Translation L1 and RMSE are normalized relative to reference trajectory RMS radius after Sim(3) alignment; rotation error reflects global orientation alignment. Runtime includes initialization and excludes depth recovery and window stitching. VGGT and VGGT-Omega process 32 evaluation frames; ViPE processes 163 – 399 frames before subsampling.
Figure 8 : Single-sample camera trajectory comparison. From left to right: COLMAP reference, VGGT, VGGT-Omega, and ViPE. All trajectories are aligned to the reference coordinate frame via Sim(3). Solid lines denote camera translation paths; schematic frustums indicate camera orientations.
τv
τa
LTX-2.3
EchoWM
HelixWorld
Pre-trained
0.10
0.01
2.16
22.58
48.69
30.14
0.10
0.05
1.96
21.85
46.52
30.34
0.10
0.10
1.55
20.87
42.15
29.30
0.10
0.20
1.16
17.57
31.65
25.91
0.10
0.30
0.90
14.91
22.94
22.80
0.20
0.01
2.41
25.92
51.53
36.92
Appendix
Table 12 : World-model sensitivity to spatial thresholds. Results use the 145 -candidate Perspective set. Pre-trained denotes the bidirectional model after progressive condition injection. Bold scores mark the best model per setting; bold thresholds mark the main-table setting. Higher is better.
τv
τa
HelixWorld (native)
+ AudioX
+ ThinkSound
+ PrismAudio
0.10
0.01
49.32
−3.94
21.72
−9.94
0.10
0.05
47.11
−5.33
21.36
−8.81
0.10
0.10
42.79
−4.84
19.04
−6.69
0.10
0.20
32.06
−4.69
14.30
−4.12
0.10
0.30
23.07
−4.82
10.47
−2.23
0.20
0.01
51.97
−3.73
19.79
−11.07
Appendix
Table 13 : Native versus post-hoc audio across spatial thresholds. All conditions use fixed HelixWorld videos from the 150 -candidate set, sharing detections and eligibility. Bold scores mark the best audio condition; bold thresholds mark the main-table setting. Higher is better.
Figure 9 : Additional visual comparisons of trajectory distillation. Each pair shows DMD only (top) and with trajectory distillation (bottom).
Figure 10 : Audio–visual synchronization during golf swings. Dashed lines mark the visible strikes in the shared video. HelixWorld’s acoustic transients coincide with both events.
Figure 11 : Audio–visual synchronization during drumming. The marked interval is a visible pause. HelixWorld’s transients subside during the pause and resume with the motion; AudioX produces a transient during the pause.
Figure 12 : Stereo audio as the fountain moves rightward. L and R denote the left and right channels. HelixWorld maintains stronger right-channel energy, consistent with the fountain’s position in the video.
Figure 13 : Stereo audio during a left-to-right vehicle pass. HelixWorld’s channel balance shifts from left to right with the car’s motion. L and R denote the two audio channels.
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Duowen Chen, Jinjin He, Gouthaman KV +2
Georgia Institute of Technology · Dolby Laboratories
Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo
Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.
Hanmo Chen, Chengcheng Liu, Tianxiao Chen +7
Xidian University · vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.