We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site
Figures & tables
Figure 1 : Overview of SSE . Given a video with degraded and unbalanced audio, SSE enhances it according to the given textual description. First, all the modalities are encoded with specific encoders, and the visual and audio features are aligned and fused. The fused and textual features are then passed to the DiT . The DiT generates enhanced audio that closely follows the original content and user guidance. Finally, the generated audio is decoded into a waveform representation with the SkipDACVAE decoder.
PANNs
PaSST
Model
FD ↓
KL ↓
FD ↓
KL ↓
IS ↑
IB ↑
Sync ↓
CLAP ↑
Input
3.55
0.35
43.39
0.26
3.16
30.49
51.80
40.56
VisAH* [ 11 ]
2.97
0.31
40.63
0.21
3.04
31.08
51.80
44.06
SSE
1.26
0.14
41.30
0.12
3.13
32.44
48.67
45.06
Table 1 : Proposed metrics calculated for DegradedMix test data. We could not evaluate VisAH-FM [ 22 ] as the model is closed-source. *: retrained with the DegradedMix data for fair comparison.
Variant
Env ↓
Mag ↓
Was ↓
KLPaSST↓
Input
6.79
23.13
1.95
25.65
VisAH* [ 11 ]
5.16
15.70
1.22
21.08
SSE
3.65
10.87
0.90
11.60
Table 2 : VGAH [ 11 ] metrics calculated for DegradedMix test data. We could not evaluate VisAH-FM [ 22 ] as the model is closed-source. *: retrained with the DegradedMix data for fair comparison.
PANNs
PaSST
Model
FD ↓
KL ↓
FD ↓
KL ↓
IS ↑
IB ↑
Sync ↓
CLAP ↑
Input
2.30
0.29
31.86
0.21
3.11
31.19
51.19
38.30
VisAH [ 11 ]
1.26
0.18
18.06
0.11
3.06
31.81
49.71
39.56
SSE
1.04
0.13
35.78
0.10
3.13
32.74
47.69
46.78
Table 3 : Proposed metrics calculated for MuddyMix [ 11 ] test data. We could not evaluate VisAH-FM [ 22 ] as the model is closed-source and the authors do not provide samples for the MuddyMix test set.
Model
Env ↓
Mag ↓
Was ↓
KLPaSST↓
Input
6.29
22.69
1.96
20.74
VisAH [ 11 ]
3.38
9.99
0.84
11.37
VisAH-FM [ 22 ]
2.74
8.28
0.63
9.70
SSE
3.46
9.86
0.88
9.61
Table 4 : VGAH [ 11 ] metrics calculated for MuddyMix [ 11 ] test data.
Comparison (A vs. B)
Wins (A–B)
Pref. rate
p -value
SSE -L vs. SSE -S
68–22
75.6%
<0.001
SSE -L vs. Original audio
72–18
80.0%
<0.001
SSE -L vs. VisAH [ 11 ]
80–10
88.9%
<0.001
SSE -S vs. Original audio
47–43
52.2%
0.752
SSE -S vs. VisAH [ 11 ]
57–33
63.3%
0.015
Original audio vs. VisAH [ 11 ]
62–28
68.9%
<0.001
Table 5 : Pairwise preference test results for UGC enhancement. Preference rate is the proportion of trials in which the first condition was preferred over the second, with 95% confidence intervals. p -values are from a two-sided binomial test against chance (50%).
PANNs
PaSST
Model
FD ↓
KL ↓
FD ↓
KL ↓
IS ↑
IB ↑
Sync ↓
CLAP ↑
Small
1.12
0.14
37.81
0.11
3.13
32.76
47.82
47.02
Base
1.06
0.13
37.28
0.10
3.12
32.64
48.84
47.10
Large
1.04
0.13
35.78
0.10
3.13
32.74
47.69
46.78
Table 6 : Ablation on different model variants. We use MuddyMix [ 11 ] data. Preferred configuration is highlighted in yellow.
PANNs
PaSST
Variant
FD ↓
KL ↓
FD ↓
KL ↓
IS ↑
IB ↑
Sync ↓
CLAP ↑
DACVAE
3.38
0.21
143.55
0.28
3.03
32.50
48.35
48.13
SkipDACVAE
1.04
0.13
35.78
0.10
3.13
32.74
47.69
46.78
Table 7 : Effect of using SkipDACVAE. We use MuddyMix [ 11 ] data. Preferred configuration is highlighted in yellow.
PANNs
PaSST
Type
FD ↓
KL ↓
FD ↓
KL ↓
IS ↑
IB ↑
Sync ↓
CLAP ↑
Original
1.24
0.16
39.15
0.12
3.13
32.58
49.97
44.82
Summarized
1.22
0.16
38.47
0.12
3.12
32.66
49.11
44.78
Audio-based
1.12
0.14
37.81
0.11
3.13
32.76
47.82
47.02
Table 8 : Ablation on different caption types. We use small-variant and MuddyMix [ 11 ] . Preferred configuration is highlighted in yellow.
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.
Automatic music mixing aims to combine multitrack recordings into a balanced and coherent musical piece. Because the content of different songs and the subjective preferences of mixing engineers jointly shape the final outcome, a practical system should deliver well-balanced mixes while allowing for controllable stylistic variation. However, most existing methods treat automatic mixing and mixing style control as separate tasks, making it difficult for a single system to produce high-quality mixes while remaining editable and style-aware. To address this limitation, this paper presents Diff2Mix, a generative automatic mixing system based on diffusion models and a differentiable mixing console. This system offers two levels of optional user control: a reference audio enables overall production style control, and the differentiable mixing console provides explicit audio effects parameters for interpretability and fine-grained optimization. We demonstrate our system's competitive performance through both objective and subjective evaluations in terms of mixing quality and control ability. We provide code and audio samples at our project page https://zys711.github.io/Diff2Mix .
Yisu Zong, Jinjie Shi, Joshua Reiss
Centre for Digital Music, Queen Mary University of London, UK
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient fine-grained supervision for real-world mixed audio and limited acoustic representations for modeling concurrent audio components. We present Dasheng AudioGen, a unified framework for generating general mixed-audio scenes from text. Dasheng AudioGen introduces structured multi-view captions, which explicitly decouple complex acoustic scenes into complementary description views, thereby enabling fine-grained control over audio layers. Furthermore, we employ a high-dimensional unified semantic-acoustic representation as the shared latent space. It injects semantic priors that facilitate cross-modal training convergence, while its high-dimensional feature space provides sufficient capacity to disentangle and fuse concurrent audio components effectively. With these designs, a simple flow-matching DiT achieves high-quality end-to-end audio scene generation. We also establish a comprehensive evaluation pipeline for audio scene generation. Experiments demonstrate that Dasheng AudioGen achieves performance approaching real-world recordings in mixed-audio categories, while remaining competitive with specialized models in single-type generation tasks. Demos are available at https://nieeim.github.io/Dasheng-AudioGen-Web/.
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
1X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China · 2MiLM Plus, Xiaomi Inc., Beijing, China