A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory. Rendering such a scene from a novel viewpoint requires that every source be available as an individual waveform, so that it can be localized in the scene and propagated to the observer before the signals are mixed. Joint audio-video generators can synthesize the video and its soundtrack, but the soundtrack is emitted as a single audio-mix in which the sources are not individually accessible. We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run. SepGen supports two complementary modes: generation and separation. In generation mode, each source caption specifies what its stem contains. A two-speaker dialogue, for example, comes out as one stem per speaker in the original turn order. In separation mode, the input audio-mix remains clean while the captions specify what to extract, so the model can decompose a recording from a free-text description. We evaluate generation on scenes synthesized from text, and separation on scenes rendered by other generators and on real recordings of speech, music, and sound effects. Given an audio-mix and captions that carry the spoken lines, SepGen outperforms language-conditioned separators, most clearly on speech, and it keeps the lead when the lines are removed from the captions. Code, checkpoints, and datasets are available at https://sepgen.github.io/
Figures & tables
Figure 1: SepGen overview and motivation. (a) Joint audio-video generation emits one soundtrack with no addressable source: the video is lifted to 4D and rerendered from a new viewpoint, but both ears hear the same mix. (b) SepGen emits one waveform per captioned source, each placed at its source in the lifted scene and rendered as spatial sound from the new viewpoint (illustration). A demonstration video is available on the project page.
Figure 2: Inference schedules in σ -space. Left to right: (a) Separation keeps the audio-mix m clean. (b) Shared-noise generation, the naive approach, denoises the stems and the audio-mix at one shared noise level. (c) SepGen’s stems attend to noisy m above σg , then its clean estimate z^0,m below it.
Sounds ( n=134 )
Speech ( n=60 )
Method
Judge ↑
J-swap ↑
CLAP ↑
C-swap ↑
WER ↓
Cosine ↓
FlowSep
3.35
1.26
0.14
0.10
1.03
0.97
AudioSep
3.23
1.27
0.14
0.11
0.91
0.87
SAM large-tv
2.96
0.41
0.05
0.03
0.87
0.95
SAM large
2.93
0.42
0.04
0.03
0.81
0.95
Ours (joint run)
3.82
2.07
0.15
0.18
0.07
0.66
Table 1: Generation benchmark. Comparison of SepGen with cascade baselines, each of which separates the same generated audio-mix given the same captions (sounds n=134 , speech n=60 ). Both SepGen rows use the generation checkpoint.
Figure 3: Stem-to-video attention in a generated frame. Each stem attends to the source named in its caption: bongo drum (left) and acoustic guitar (right); warmer is higher.
Sounds ( n=104 )
Speech ( n=45 )
Method
Judge ↑
J-swap ↑
CLAP ↑
C-swap ↑
WER ↓
Cosine ↓
AudioSep
3.46
1.60
0.19
0.16
0.93 / 1.03
0.89 / 0.86
FlowSep
3.21
1.33
0.15
0.10
0.98 / 0.99
0.94 / 0.94
SAM large
3.14
0.50
0.07
0.04
0.74 / 0.73
0.95 / 0.94
SAM large-tv
3.06
0.35
0.07
0.04
0.81 / 0.79
0.96 / 0.94
SA × SAM3
3.07
0.27
0.08
0.03
0.87 / –
0.94 / –
Table 2: Separation on the Veo benchmark (149 scenes rendered by Veo 3.1 Lite). Speech cells report results with and without the quoted lines.
Method
Judge ↑
J-swap ↑
CLAP ↑
C-swap ↑
SAM large-tv
3.12
0.78
0.13
0.03
AudioSep
2.81
0.82
0.20
0.07
SAM large
2.75
0.63
0.10
0.01
FlowSep
2.64
0.51
0.16
0.03
SepGen (separation checkpoint)
3.70
1.42
0.21
0.11
SepGen (generation checkpoint)
3.70
1.52
0.21
0.13
Table 3: Separation on SAM-Audio-bench . All methods are queried with identical captions.
Figure 4: User study. Forced choice with a tie option, against SAM Audio (a, left) on ten clips from Tables 2 and 3 , 48 listeners, and AudioSep (b, right) on ten generation clips of Table 1 , 51 listeners. Bars give the share of the n judgements.
Sounds
Speech
Component
Condition
Judge ↑
J-swap ↑
CLAP ↑
C-swap ↑
Sib-sim ↓
WER ↓
Cosine ↓
Our configuration
4.07
1.99
0.19
0.18
0.01
0.09
0.66
Jointness
Independent passes
2.43 ***
0.21 ***
0.16 *
0.04 ***
0.13 ***
1.05 ***
0.95 ***
Video
Blank
4.09
2.01
0.19
0.18
0.01
0.09
0.67
Unrelated
4.08
2.01
0.19
0.17
0.01
0.09
0.68
None
4.06
1.99
0.19
0.17
0.01
0.13
0.67
Table 4: Ablations on the LTX-2.3 benchmark (separation checkpoint, seed 0). Each block modifies a single component of our configuration. Stars mark Holm-corrected paired Wilcoxon significance against our configuration (* p<0.05 , ** p<0.01 , *** p<0.001 ).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Cross-Stem Attention Guidance. (a) Each stem attends to its own caption in the positive branch (solid) and its sibling’s in the negative branch (dashed). The audio-mix attends to c in both. (b) NAG moves Z+ away from Z− to obtain ZCSAG .
Sounds
Speech
Guidance
Judge ↑
J-swap ↑
CLAP ↑
C-swap ↑
Sib-sim ↓
WER ↓
Cosine ↓
λ=2 (ours)
4.07
1.99
0.19
0.18
0.01
0.09
0.66
Off
3.92
1.80
0.18
0.15 *
0.03
0.08
0.68
λ=3
3.94 *
1.77 *
0.18
0.18
0.01
0.16
0.66
λ=5
3.87 *
1.74
0.18
0.16
0.01
0.17
0.66
Appendix
Table S1: Cross-Stem Attention Guidance ablation on the LTX-2.3 benchmark (separation checkpoint, seed 0); stars mark Holm-corrected paired Wilcoxon significance against λ=2 , as in Table 4 of the main paper.
Method
Judge (published)
Δ
FlowSep
3.49
+0.08
AudioSep
3.27
+0.15
SAM large-tv
2.84
−0.09
SAM large
2.75
−0.08
SepGen (second round trip)
3.82
0.00
Appendix
Table S2: Audio VAE requantization control (sounds, n=134 , seed 0). The stems of each method are passed once more through the audio VAE; Δ denotes the requantized minus the published score.
Study
Cut
n
SepGen
Tie
Baseline
Sources
Separation
Veo
336
0.51 [0.40, 0.58]
0.19
0.30 [0.19, 0.42]
3/0/4
SAM-Audio-bench
384
0.64 [0.45, 0.83]
0.16
0.20 [0.09, 0.34]
6/0/2
All
720
0.58 [0.47, 0.70]
0.17
0.25 [0.16, 0.34]
9/0/6
Generation
Sounds
612
0.49 [0.35, 0.63]
0.13
0.38 [0.24, 0.51]
6/1/5
Speech
357
0.87 [0.80, 0.95]
0.03
0.10 [0.04, 0.17]
7/0/0
All
969
0.63 [0.48, 0.78]
0.09
0.28 [0.15, 0.40]
13/1/5
Appendix
Table S3: User-study preference shares with 95% bootstrap intervals, resampling listeners and clips. The Sources column reports the number of sources on which SepGen won, tied, or lost by majority vote.
Figure S2: Generated stems. Top: singing and drums. Bottom: two turn-taking speakers. Each row shows the generated video, audio-mix, and caption-addressed stems from one sampling run.
Sounds
Speech
Method
Judge ↑
J-swap ↑
WER ↓
Cosine ↓
SepGen (joint run)
3.82
2.08
0.11
0.67
Independent passes
2.27
0.65
0.69
0.86
FlowSep on the joint audio-mix
3.49
1.39
1.02
0.96
Appendix
Table S6: Effect of joint sampling on the generation benchmark (seed 0): a single joint run compared with one single-caption run per source.
Method
Captions
Judge ↑
J-swap ↑
CLAP ↑
C-swap ↑
SAM large-tv
ours
3.12±0.08
0.78±0.14
0.13±0.01
0.03±0.01
native
3.32±0.11
0.98±0.12
0.14±0.01
0.04±0.01
AudioSep
ours
2.81±0.00
0.82±0.00
0.20±0.00
0.07±0.00
native
2.38±0.00
0.51±0.00
0.20±0.00
0.03±0.00
SAM large
ours
2.75±0.15
0.63±0.06
0.10±0.01
0.01±0.01
native
2.70±0.13
0.58±0.09
0.10±0.01
0.01±0.01
Appendix
Table S10: SAM-Audio-bench (172 clips). Each baseline is queried both with our captions ( ours ) and with the native benchmark text ( native ); rankings are computed over the ours rows.
Audio-mix clamped to
Judge ↑
Word error
Own-line ↓
Donor-line
Observed scene
3.80
0.09
0.99
Cross-clip donor
2.97
1.16
0.41
Digital silence
1.10
1.00
–
Appendix
Table S12: Audio-mix ablation on the LTX-2.3 clip set (separation checkpoint, 64 clips), in which the observed audio-mix is replaced by silence or by the audio-mix of another clip.
Method
Envelope r↑
LSD ↓
SepGen (separation checkpoint)
0.82
6.3
AudioSep
0.76
11.0
FlowSep
0.65
13.1
SAM large-tv
0.47
17.6
SAM large
0.39
18.9
SAM large-tv, best of eight
0.78
10.3
Appendix
Table S13: Reference-stem evaluation on 129 held-out scenes with ground-truth stems (258 stems, seed 0), measured by envelope correlation and log-spectral distance. Best-of-eight rows and controls are not ranked.
Stem vs. round trip
SI-SDR (dB), restoring original
Method
SI-SDR (dB)
envelope r
LSD (dB)
phase
magnitudes
FlowSep
−25.2
0.99
2.8
+15.3
−25.0
AudioSep
−13.5
0.99
2.6
+15.7
−13.2
SAM large-tv
−26.1
0.78
2.0
+6.5
−24.5
SAM large
−25.1
0.81
2.1
+7.5
−23.8
SepGen (2nd round trip)
−10.4
0.96
2.2
+15.8
−9.7
Appendix
Table S14: Audio VAE decoder round trip (194 clips, seed 0). Top: the stems of each method evaluated against their own round trip. Bottom: the summed stems evaluated against the audio-mix.
Figure S3: Same-pass placement readout on two scenes. Left: instrument couple (djembe / cello). Right: turn-taking dialogue (tattoo artist / client). Each column shows the observed video and mix, the two captions, the H×W×T attention volume (16 latent frames), the two source waveforms, and the pooled geometric-median + Kalman tracks. Voxels and tracks are drawn only on latent bins where that source waveform is active; quiet bins are omitted. Colors match the two captions.
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient fine-grained supervision for real-world mixed audio and limited acoustic representations for modeling concurrent audio components. We present Dasheng AudioGen, a unified framework for generating general mixed-audio scenes from text. Dasheng AudioGen introduces structured multi-view captions, which explicitly decouple complex acoustic scenes into complementary description views, thereby enabling fine-grained control over audio layers. Furthermore, we employ a high-dimensional unified semantic-acoustic representation as the shared latent space. It injects semantic priors that facilitate cross-modal training convergence, while its high-dimensional feature space provides sufficient capacity to disentangle and fuse concurrent audio components effectively. With these designs, a simple flow-matching DiT achieves high-quality end-to-end audio scene generation. We also establish a comprehensive evaluation pipeline for audio scene generation. Experiments demonstrate that Dasheng AudioGen achieves performance approaching real-world recordings in mixed-audio categories, while remaining competitive with specialized models in single-type generation tasks. Demos are available at https://nieeim.github.io/Dasheng-AudioGen-Web/.
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
1X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China · 2MiLM Plus, Xiaomi Inc., Beijing, China
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.