Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work, we use activation patching to demonstrate that recent audio diffusion architectures exhibit a semantic bottleneck, where a small, shared subset of consecutive attention layers controls distinct musical concepts, such as the presence of specific instruments, vocals, or genres. Building on this, we systematically evaluate a broad spectrum of steering paradigms, comparing activation steering against prompt-level, score-space, and weight-space interventions, analyzing the interaction between the steering mechanism and the intervention site. Our new benchmark, supported by an extensive user study, demonstrates that localized activation steering establishes a new state-of-the-art in audio concept modulation.
Figures & tables
Figure 1 : We study localized activation steering in Audio Diffusion Models. By localizing functional layers, we enable precise modulation of audio concepts through activation steering.
Figure 2 : Layer localization via Activation Patching. For a target concept, we perform a clean run (a) with a prompt referring to the concept and save the attention key ( K ) and value ( V ) activations. In a corrupted run (b), we infer with a counterfactual prompt. Next (c,d), we run a corrupted run, substituting the l -th layer key and value activations with the saved ones from the clean run. If (d) patching inside l produces audio with the target concept, we identify it as a functional layer.
Figure 3 : Average impact scores I(l) of cross-attention layers (denoted with color intensity in □ ) within text-to-music DMs. Singular layers control different concepts: vocal, tempo, mood, instruments, and genres across diverse audio diffusion architectures. Plots for all concepts are provided in Appendix F .
Method
AUC ( ↑ )
Smoothness ( ↓ )
Audio Quality ( ↑ )
MuQ
CLAP
MuQ
CLAP
Prompt-level interventions
PCI ( Görgün et al. [26] )
0.084±.019
0.049±.010
0.185±.064
0.223±.066
6.803±.016
Text Embeddings ( Ezra et al. [17] )
0.014±.007
0.011±.004
0.439±.224
0.681±.277
6.690±.027
Token Embeddings ( Baumann et al. [7] )
0.035±.011
0.024±.006
0.267±.063
0.211±.058
6.776±.023
Score-space steering
Table 1 : Audio concept modulation with Ace-Step. Localized activation steering outperfroms prompt-, score-, and weight-level methods. Results are averaged over 9 musical concepts, steered in both positive and negative directions (see App. M for details). Best and second results are highlighted.
Figure 4 : Human evaluation of audio steering methods. We report mean Likert rating ( 1 – 5 ) for each method in the standard (all layers) and localized configurations across the three questions.
Figure 5 : Localized activation steering outperforms the global approach. For the same audio preservation level, intervening in functional layers offers higher gain in the alignment. Additionally, omitting functional layers for steering often yields no conceptual gain.
Method
AUC ( ↑ )
Smoothness ( ↓ )
Audio Quality ( ↑ )
MuQ
CLAP
MuQ
CLAP
PCI ( Görgün et al. [26] )
0.084±.019
0.049±.010
0.185±.064
0.223±.066
6.803±.016
Δ (localization)
−19%
−22%
−50%
+21%
−0.2%
Text Embeddings ( Ezra et al. [17] )
0.014±.007
0.011±.004
0.439±.224
0.681±.277
6.690±.027
Δ (localization)
+62%
+31%
+25%
+48%
+1.4%
Token Embeddings ( Baumann et al. [7] )
0.035±.011
0.024±.006
0.267±.063
0.211±.058
6.776±.023
Table 2 : Effect of layer localization on each steering method. We compare the methods’ global variant (all layers steering) against the localized one. Gain from localization is expressed as a % change of the mean: improvement or degradation of steering quality, faded when the absolute gain is below the standard error of the difference. We provide full results in App. M.1 .
Figure 6 : Human preference for localized vs. global setups of steering methods. We compare user study ratings of both configurations for the same method. Bars present pairs in which the localized variant is preferred (green), the two are tied (grey), or the global steering is preferred (blue).
Method
Steering Two Concepts
Steering Three Concepts
P + V
P + FV
A + FV
P + FV
J + T
P + V + J
A + FV + M
P + V + T
A + FV + M
AUSteer
+0.057
+0.001
+0.022
+0.013
+0.090
+0.062
+0.005
+0.048
+0.045
CAA
+0.048
+0.011
+0.028
+0.008
+0.068
+0.047
+0.026
+0.029
+0.032
AUSteer (loc.)
+0.099
+0.018
+0.029
+0.039
+0.086
+0.088
+0.021
+0.069
+0.051
CAA (loc.)
+0.094
+0.034
+0.052
+0.027
+0.103
+0.081
+0.035
+0.061
+0.039
SAE (loc.)
+0.114
+0.027
+0.041
+0.034
+0.110
+0.100
+0.018
+0.081
+0.063
Table 3: Multi-concept steering performance measured with average Area Under LPAPS-MUQ Curve. Best result per combination of concepts is bolded . Single column denotes steering on combinations of concepts, where P = Piano, V = Violin, FV = Female Vocal, A = Acoustic Guitar, J = Jazz Music, M = Brighter Mood, FV = Male Vocal, T = Slower Tempo, and M = Darker Mood. In App. O , we provide more details regarding the experiment.
Appendix figures & tables51 assets
Supplementary material from the paper’s appendix.
Appendix
Concept
Original Prompt Pc
Counterfactual Prompt Pc~
Vocal Gender
"The low quality recording features a ballad song that contains sustained strings, mellow piano melody, and soft female vocal singing over it. It sounds sad and soulful."
"The low quality recording features a ballad song that contains sustained strings, mellow piano melody, and soft male vocal singing over it. It sounds sad and soulful."
Tempo
"This is a country music piece. There is a fiddle playing the main melody. The acoustic guitar and electric guitar are playing gently. The song has a slow tempo. The atmosphere is sentimental."
"This is a country music piece. There is a fiddle playing the main melody. The acoustic guitar and electric guitar are playing gently. The song has a fast tempo. The atmosphere is sentimental."
Mood
"This is a Hindustani classical music piece. There is a harmonium playing the main tune. A bansuri joins in to play, supporting a melody. The rhythmic background consists of tabla percussion and electronic drums. The atmosphere is joyful ."
"This is a Hindustani classical music piece. There is a harmonium playing the main tune. A bansuri joins in to play, supporting a melody. The rhythmic background consists of tabla percussion and electronic drums. The atmosphere is sorrowful ."
Instrument
"This folk song features a male voice singing the main melody in an emotional mood. This is accompanied by an accordion playing fills in the background. A violin plays a droning melody."
"This folk song features a male voice singing the main melody in an emotional mood. This is accompanied by an accordion playing fills in the background. A trumpet plays a droning melody."
Genre
"The low quality recording features a reggae /dub song that consists of a flat male vocal singing over punchy 808 bass, punchy snare, shimmering hi hats and groovy piano chords. It sounds energetic, groovy and the recording is noisy and in mono."
"The low quality recording features a metal /dub song that consists of a flat male vocal singing over punchy 808 bass, punchy snare, shimmering hi hats and groovy piano chords. It sounds energetic, groovy and the recording is noisy and in mono."
Appendix
Table 4: Examples of counterfactual prompt pairs. For each concept category, we show an original prompt from MusicCaps and its counterfactual version with the target concept replaced. Modified terms are highlighted in bold .
sad → happy lonely → connected emotional → lighthearted
Appendix
Table 5: Keywords for dataset construction. For each concept, we list the examples of keywords proposed by an LLM (GPT-4o) to filter captions from MusicCaps (for a given concept, we select prompts containing these terms) and the replacement keywords used to generate counterfactual prompts.
Figure 7 : Impact of every cross-attention layer on every concept in AudioLDM2 [ 37 ] . Each cell color intensity denotes layer l impact on concept c as I(l,c) ( Eq. 2 ). Last row denotes the aggregated impact I(l) of layer l averaged over all concepts. The final selected layers (where I(l)≥τ=0.10 ) are outlined in orange.
Figure 8 : Impact of every cross-attention layer on every concept in Stable Audio Open [ 16 ] . Each cell color intensity denotes layer l impact on concept c as I(l,c) ( Eq. 2 ). Last row denotes the aggregated impact I(l) of layer l averaged over all concepts. The final selected layers (where I(l)≥τ=0.10 ) are outlined in orange.
Figure 9 : Impact of every cross-attention layer on every concept in Ace-Step [ 25 ] . Each cell color intensity denotes layer l impact on concept c as I(l,c) ( Eq. 2 ). Last row denotes the aggregated impact I(l) of layer l averaged over all concepts. The final selected layers (where I(l)≥τ=0.10 ) are outlined in orange.
Prompts / concept
Budget %
J (top-1)
J (top-2)
J (top-3)
J (top-5)
4
2%
0.54
0.82
0.66
0.45
12
5%
0.65
0.98
0.81
0.56
17
8%
0.67
0.99
0.90
0.63
23
10%
0.60
0.99
0.92
0.67
34
15%
0.65
1.00
0.97
0.74
46
20%
0.69
1.00
0.99
0.80
Appendix
Table 6 : Impact of number of prompts on ACE-Step localization results. J (top-j) denotes Jaccard similarity between sets of j highest scoring layers selected by full-budget localization and a random subset of prompts. Results are averaged across 200 random subsamples.
τ
AudioLDM2
Stable Audio Open
Ace-Step
0.05
45, 46, 49, 50, 51
5, 6, 12, 13, 14, 19
6, 7, 8
0.075
45, 46, 49, 50, 51
5, 12, 13, 14
6, 7, 8
0.10
45, 46, 51
12, 13, 14
7, 8
0.15
45, 51
12, 13
7, 8
0.20
45, 51
12, 13
7, 8
0.30
45, 51
13
8
Appendix
Table 7: Selected semantic bottleneck in Audio Diffusion Models as a function of selection threshold τ . Setting used in our experiments is bolded .
Concept
Positive Prompt Pc
Negative Prompt Pc~
Piano
“{base}, with piano”
“{base}”
Violin
“{base}, with violin”
“{base}”
Guitar Type
“{base}, with acoustic guitar”
“{base}, with electric guitar”
Mood
“happy song, {base}”
“sad song, {base}”
Tempo
“fast song, {base}”
“slow song, {base}”
Vocal Gender
“{base}, with female vocal”
“{base}, with male vocal”
Appendix
Table 8: Contrastive prompt templates for steering. The {base} placeholder is filled with diverse musical descriptions (e.g., “a song”, “a jazz piece”, “electronic music”).
Concept
Alignment query
Piano
“a piano song”
Violin
“a song with violin”
Guitar Type
“a song with acoustic guitar”
Mood
“a cheerful track”
Tempo
“a fast track”
Vocal Gender
“female vocal singing”
Appendix
Table 9 : Audio-text alignment queries for audio-text alignment measurement. These prompts are used to compute similarity scores between generated audio and target concepts.
Figure 10 : Text-audio alignment models track the presence of concepts in generated audios. We compare alignment metrics from our evaluation benchmark with external acoustic evaluators on mood , tempo , and vocal gender concepts. We plot the per- α averages of MuQ, CLAP, and the external evaluator, each normalized to [0,1] . As shown, CLAP and MuQ track real audio properties. In legend, we report the Pearson correlation ρ between each text-based metric and the evaluator.
Method
LPAPS ( ↑ )
Harmony ( ↑ )
Rhythm ( ↑ )
Melody ( ↑ )
Structure ( ↑ )
MuQ
CLAP
MuQ
CLAP
MuQ
CLAP
MuQ
CLAP
MuQ
CLAP
PCI
0.021
0.012
0.026
0.015
0.020
0.013
0.015
0.009
0.020
0.012
PCI (loc.)
0.017
0.009
0.022
0.012
0.015
0.009
0.011
0.006
0.015
0.009
Text Embeddings
0.004
0.003
0.005
0.003
0.003
0.003
0.002
0.002
0.003
0.003
Text Embeddings (loc.)
0.006
0.004
0.007
0.004
0.004
0.003
0.004
0.002
0.005
0.003
Token Embeddings
0.009
0.006
0.010
0.007
0.008
0.006
0.006
0.005
0.008
0.006
Appendix
Table 10 : Localized steering offers best steering performance across diverse preservation metrics. We report AUC-MuQ and AUC-CLAP under LPAPS and four MIR preservation axes, measuring harmony, rhythm, melody, and structure. Results are averaged over the nine benchmark concepts (except tempo on the Rhythm axis) and across both steering directions. Best and second results are highlighted.
Figure 11 : Per-concept steering methods comparison across different preservation metrics: Harmony, Rhythm, Melody, Structure. We report preservation–alignment AUC per concept, averaged over both alignment metrics: MuQ and CLAP. Hatched bars represent localized variants of the methods. Localized activation steering (CAA, AUSteer, SAE) leads on nearly every concept and axis.
α/αmax
CAA
CAA (loc.)
AUSteer
AUSteer (loc.)
SAE (loc.)
0.0
0.000
0.000
0.000
0.000
0.000
0.2
−0.028
+0.024
+0.008
+0.030
+0.046
0.4
−0.045
+0.041
−0.043
+0.043
+0.067
0.6
−0.051
+0.049
−0.090
+0.015
+0.045
0.8
−0.055
+0.059
−0.137
−0.001
+0.034
1.0
−0.047
+0.096
−0.154
−0.010
+0.030
Appendix
Table 11 : Change in Production Quality (PQ) under increasing strength in activation steering methods. We report Δ PQ relative to the unsteered baseline ( PQ=7.40 ) at increasing steering strengths. Localized variants show little-to-no degradation at any strength, whereas global methods decline as steering strength ( α ) increases.
Learning rate
Train guidance
Training iters
Value
1e-4 ⋆
5e-4
5e-3
3
5
7 ⋆
500 ⋆
1000
1500
Avg AUC
0.102
0.101
0.083
0.104
0.103
0.102
0.102
0.101
0.101
Appendix
Table 12 : Hyperparameter impact on audio concept modulation with Concept Sliders. We report preservation-alignment AUC (averaged over piano and vocal concepts) on the held-out dataset. The setting used in the experiments is marked ⋆ , the best per axis is in bold .
Rank
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
Avg
4
0.061
0.039
0.038
0.104
0.080
0.162
0.069
0.076
0.088
0.080
8
0.061
0.036
0.038
0.102
0.082
0.155
0.071
0.082
0.091
0.080
16
0.061
0.036
0.043
0.103
0.080
0.165
0.075
0.082
0.090
0.082
Appendix
Table 13 : Per-concept LoRA-rank sweep for Concept Sliders. We report preservation–alignment AUC. The per-concept highest scoring parameter (used in the benchmark) is bolded .
Top- s
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
Avg
All-layer steering
256
0.049
0.034
0.035
0.034
0.038
0.024
0.052
0.087
0.097
0.050
1024
0.048
0.046
0.042
0.037
0.030
0.069
0.068
0.108
0.135
0.065
2048
0.054
0.039
0.048
0.039
0.022
0.082
0.075
0.118
0.143
0.069
4096
0.040
0.030
0.055
0.063
0.034
0.106
0.069
0.109
0.151
0.073
8192
0.024
0.028
0.046
0.065
0.029
0.124
0.051
0.105
0.115
0.065
Appendix
Table 14 : Per-concept top- s sweep for AUSteer. We report preservation–alignment AUC (mean of MuQ-T and CLAP over both steering directions) on the held-out dataset. The best value in each column is bolded .
Method
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
Avg
add. (eq. 9 )
0.107
0.031
0.042
0.087
0.102
0.219
0.109
0.191
0.211
0.122
mult. (eq. 8 )
0.028
−0.005
−0.009
0.028
0.003
−0.008
0.009
0.028
−0.013
0.007
Appendix
Table 15 : Impact of steering vector construction on AUSteer performance. We report preservation–alignment AUC on the held-out set over all concepts and highlight the best method.
Layer
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
5
9.6
0.3
1.3
0.3
17.5
11.1
1.8
0.2
17.7
6
5.3
0.0
0.4
0.0
8.4
2.7
2.2
0.0
0.2
7
0.6
0.0
17.6
17.6
1.4
3.0
1.6
8.8
1.6
8
1.8
12.0
30.1
37.6
14.9
17.3
62.1
53.5
71.8
9
0.0
0.7
1.1
0.0
0.2
15.2
0.2
2.5
0.4
12
0.4
0.0
0.0
5.8
0.6
0.1
0.1
0.0
0.1
Appendix
Table 16 : Ace-Step layers allocated with AUSteer’s rank-and-select localization. We report the percentage of activation dimensions in attention layers, averaged across diffusion steps, for all nine concepts. Layers with at least 1% allocation to at least one concept are listed, and the functional layers ( {7,8} ) we localize are highlighted.
Steered CFG passes
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
Avg
cond
0.104
0.035
0.059
0.075
0.105
0.239
0.117
0.192
0.191
0.124
cond + null
0.110
0.040
0.058
0.089
0.110
0.230
0.114
0.203
0.177
0.126
Appendix
Table 17 : Impact of model’s classifier-free guidance (CFG) prediction on concept steering with CAA. We report preservation-alignment AUC over concepts. Steering separately both conditional and unconditional ( cond + null ) predictions does not yield significantly different results than steering only the conditional prediction ( cond ).
ACE-Step Layer
FVU ( ↓ )
dead% ( ↓ )
fire% ( ↓ )
7
0.216
0.37%
19.4%
8
0.243
0.25%
20.1%
Appendix
Table 18 : Training quality metrics for our BatchTopK SAEs. Lower FVU indicates better SAE reconstruction, lower dead% indicates a smaller fraction of features that never activate; lower fire% indicates fewer features that activate every batch (global).
Figure 12 : Feature-density histograms for Sparse Autoencoders trained Ace-Step semantic bottleneck. Features in SAE latent space are considered dead (never used) when their density <10−7 , while high-frequency features (activating at least once per batch) are the ones with density >10−2 . We consider features in between to be prospective for steering.
Figure 13 : Impact of SAE feature count ( k ) on each concept modulation performance. We find that optimal configuration of k for layers (attn7,attn8) highly depends on given concept. In our experiments, we use the best-scoring pair for each concept, which is highlighted.
Weighting
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
Avg
unit ⋆
0.124
0.031
0.065
0.126
0.117
0.233
0.123
0.236
0.174
0.137
TF
0.110
0.040
0.034
0.129
0.074
0.239
0.137
0.173
0.143
0.120
TF–IDF
0.095
0.032
0.041
0.112
0.092
0.220
0.143
0.169
0.160
0.118
TF, Σ=k
0.128
0.040
0.052
0.138
0.090
0.247
0.167
0.225
0.143
0.137
TF–IDF, Σ=k
0.113
0.033
0.053
0.104
0.108
0.215
0.160
0.225
0.165
0.131
Appendix
Table 19 : Impact of SAE’s features weighting on steering performance. We report AUC on the held-out set. While weighting decoder columns by raw scores yields suboptimal results, rescaling scores (Σ=k) yields performance competitive with our ( ⋆ ) approach.
Negative SV Construction
Piano
Mood
Tempo
Voc. gen.
Violin
Voc. style
Guitar
Rock
Electr.
Avg
negated positive dir. vector ⋆
0.065
0.083
0.063
0.046
0.070
0.075
0.081
0.091
0.123
0.077
TF–IDF on negative prompts
0.008
0.033
0.058
0.035
0.075
0.045
0.016
0.028
0.039
0.037
Appendix
Table 20 : Constructing negative direction (concept-removal) steering vectors with SAE. We report AUC on the held-out dataset. Our ( ⋆ ) approach of applying a positive (concept-amplifying) vector with negative steering strength performs better than identifying a negative-feature set.
Figure 14 : Print screen of the user study interface
Method (loc.)
Casual listeners ( ≤2 )
Trained musicians ( ≥3 )
Δ
PCI
3.20
2.80
−0.40
Text Embeddings
2.69
2.32
−0.37
Token Embeddins
2.64
2.58
−0.06
FreeSliders
3.38
2.71
−0.67
CAA
3.23
3.40
+0.17
AUSteer
3.10
3.42
+0.32
Appendix
Table 21: Mean Seamless-Edit rating per method (Localized variant), split by self-reported musical skill. Casual listener = skill ≤2 ; trained musician = skill ≥3 . Activation-space methods are preferred more strongly by trained listeners; score- and prompt-level methods less so.
Figure 15 : Layer selection in activation steering is crucial for audio concept modulation. Steering the activations in localized layers outperforms steering all of the layers, while steering all layers except the localized ones leads to drastic degradation.
Concept
Layers
AUC (MuQ) ↑
AUC (CLAP) ↑
Smoothness (MuQ) ↓
Smoothness (CLAP) ↓
Piano
all
0.331
0.175
0.081
0.098
localized
0.930
0.390
0.111
0.064
ablated
0.099
0.099
0.176
0.171
Tempo
all
0.236
0.008
0.125
0.222
localized
0.300
0.067
0.065
0.073
ablated
0.201
-0.007
0.135
0.605
Appendix
Table 22 : Per-concept steering metrics for AudioLDM2 with CAA. Higher is better for AUC, lower for Smoothness. The best per row is bolded .
Concept
Layers
AUC (MuQ) ↑
AUC (CLAP) ↑
Smoothness (MuQ) ↓
Smoothness (CLAP) ↓
Piano
all
0.792
0.228
0.093
0.105
localized
0.742
0.220
0.128
0.154
ablated
0.383
0.094
0.086
0.129
Tempo
all
0.268
0.108
0.110
0.100
localized
0.323
0.216
0.113
0.085
ablated
0.183
0.031
0.096
0.202
Appendix
Table 23 : Per-concept steering metrics for Stable Audio Open with CAA. Higher is better for AUC, lower for Smoothness. The best per row is bolded .
Concept
Layers
AUC (MuQ) ↑
AUC (CLAP) ↑
Smoothness (MuQ) ↓
Smoothness (CLAP) ↓
Piano
all
0.060
0.026
0.038
0.063
localized
0.141
0.061
0.033
0.032
ablated
-0.008
-0.010
0.488
–
Tempo
all
0.095
0.022
0.032
0.075
localized
0.111
0.023
0.029
0.042
ablated
0.059
0.018
0.034
0.076
Appendix
Table 24 : Per-concept steering metrics for ACE-Step with CAA. Higher is better for AUC, lower for Smoothness. The best per row is bolded .
Method
AUC ( ↑ )
Smoothness ( ↓ )
Audio Quality ( ↑ )
MuQ
CLAP
MuQ
CLAP
PCI ( Görgün et al. [26] )
0.084±.019
0.049±.010
0.185±.064
0.223±.066
6.803±.016
+ localized
0.068±.016
0.038±.008
0.277±.143
0.122±.017
6.789±.014
Δ
−19%
−22%
−50%
+21%
−0.2%
Text Embeddings ( Ezra et al. [17] )
0.014±.007
0.011±.004
0.439±.224
0.681±.277
6.690±.027
+ localized
0.023±.006
0.014±.004
0.334±.123
0.366±.140
6.780±.014
Appendix
Table 25 : Effect of layer localization on each steering method. For every method, we compare its standard variant (steering all diffusion model layers) against our localized variant (applying the intervention only at layers {7,8}). Mean ± standard error pooled across 9 musical concepts and both steering directions (see Tables 26 , 31 , 32 , 27 , 28 , 29 , 30 , 33 and 34 in Appendix). The Δ row reports the relative gain from localization, expressed as a % change of the Standard mean (sign chosen so that “ + ” always denotes improvement): green when localization helps and red when it hurts, faded when the absolute gain is below the standard error of the difference.
Method ( piano )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.104
0.004
0.054
0.044
0.001
0.022
0.066
0.418
0.242
0.077
0.501
0.289
6.814
PCI (loc.)
0.081
0.007
0.044
0.027
0.010
0.018
0.080
0.547
0.313
0.094
0.270
0.182
6.809
Text Embeddings
0.056
0.005
0.030
0.027
0.022
0.024
0.065
0.415
0.240
0.073
0.051
0.062
6.803
Text Embeddings (loc.)
0.060
0.015
0.038
0.020
0.019
0.019
0.069
0.385
0.227
0.075
0.151
0.113
6.796
Token Embeddings
0.127
0.040
0.084
0.049
0.021
0.035
0.084
0.085
0.084
0.078
0.082
0.080
6.793
Appendix
Table 26 : Evaluating steering piano concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( add piano ) and negative ( remove piano ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( tempo )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.052
0.075
0.063
0.008
0.027
0.017
0.081
0.055
0.068
0.190
0.076
0.133
6.746
PCI (loc.)
0.047
0.054
0.051
0.007
0.019
0.013
0.081
0.059
0.070
0.204
0.070
0.137
6.742
Text Embeddings
0.015
0.011
0.013
−0.016
0.021
0.002
0.234
0.115
0.175
2.731
0.072
1.402
6.678
Text Embeddings (loc.)
0.016
0.015
0.016
−0.001
−0.001
−0.001
0.208
0.119
0.163
0.404
0.684
0.544
6.782
Token Embeddings
0.012
0.025
0.018
0.011
0.020
0.015
0.166
0.150
0.158
0.082
0.088
0.085
6.741
Appendix
Table 27 : Evaluating steering tempo concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( fast song ) and negative ( slow song ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( mood )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.028
0.027
0.027
0.020
0.049
0.034
0.102
0.097
0.099
0.103
0.104
0.103
6.766
PCI (loc.)
0.008
0.031
0.020
0.008
0.044
0.026
0.178
0.077
0.127
0.103
0.069
0.086
6.748
Text Embeddings
0.010
−0.001
0.004
−0.002
0.019
0.008
0.281
0.181
0.231
0.573
0.062
0.318
6.693
Text Embeddings (loc.)
0.007
0.009
0.008
0.000
0.027
0.014
0.178
0.184
0.181
0.407
0.084
0.246
6.764
Token Embeddings
0.002
0.001
0.001
−0.005
0.026
0.010
0.462
0.265
0.364
0.938
0.116
0.527
6.710
Appendix
Table 28 : Evaluating steering mood concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( happy song ) and negative ( sad song ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( vocal gender )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.045
0.010
0.027
0.132
0.058
0.095
0.088
0.249
0.168
0.053
0.109
0.081
6.882
PCI (loc.)
0.021
0.015
0.018
0.101
0.023
0.062
0.123
0.316
0.219
0.046
0.105
0.076
6.845
Text Embeddings
−0.075
0.041
−0.017
−0.017
0.028
0.006
−
0.096
−
4.537
0.074
2.306
6.693
Text Embeddings (loc.)
−0.008
0.001
−0.004
0.025
0.007
0.016
0.187
0.563
0.375
0.088
0.119
0.103
6.821
Token Embeddings
−0.005
−0.008
−0.007
0.004
0.010
0.007
0.597
0.454
0.525
0.160
0.241
0.201
6.882
Appendix
Table 29 : Evaluating steering vocal gender concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( female vocal ) and negative ( male vocal ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( vocal style )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.155
−0.010
0.073
0.056
−0.001
0.028
0.068
1.277
0.672
0.074
1.250
0.662
6.823
PCI (loc.)
0.136
−0.010
0.063
0.053
−0.011
0.021
0.069
2.739
1.404
0.091
−
−
6.800
Text Embeddings
−0.005
−0.014
−0.010
−0.026
0.006
−0.010
1.029
2.983
2.006
0.630
0.122
0.376
6.674
Text Embeddings (loc.)
0.036
−0.002
0.017
0.024
−0.010
0.007
0.077
2.508
1.293
0.246
2.549
1.398
6.811
Token Embeddings
0.045
0.007
0.026
0.032
0.004
0.018
0.089
0.190
0.139
0.081
0.197
0.139
6.877
Appendix
Table 30 : Evaluating steering vocal style concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( rap vocal ) and negative ( sing vocal ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( violin )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.148
0.018
0.083
0.065
0.011
0.038
0.076
0.239
0.158
0.075
0.186
0.130
6.828
PCI (loc.)
0.127
0.019
0.073
0.051
0.014
0.033
0.064
0.146
0.105
0.076
0.165
0.120
6.833
Text Embeddings
0.088
−0.006
0.041
0.045
0.007
0.026
0.055
0.472
0.263
0.045
0.105
0.075
6.824
Text Embeddings (loc.)
0.089
0.004
0.047
0.033
0.016
0.025
0.051
0.473
0.262
0.057
0.126
0.092
6.822
Token Embeddings
0.146
0.027
0.087
0.062
0.018
0.040
0.037
0.175
0.106
0.041
0.104
0.072
6.817
Appendix
Table 31 : Evaluating steering violin concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( add violin ) and negative ( remove violin ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( guitar type )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.146
0.012
0.079
0.073
0.001
0.037
0.059
0.139
0.099
0.059
0.762
0.410
6.730
PCI (loc.)
0.117
0.013
0.065
0.049
0.015
0.032
0.056
0.139
0.098
0.067
0.329
0.198
6.723
Text Embeddings
0.010
0.007
0.009
−0.029
0.026
−0.002
0.139
0.305
0.222
−
0.068
−
6.569
Text Embeddings (loc.)
0.005
0.005
0.005
−0.020
0.019
−0.001
0.193
0.300
0.247
0.680
0.235
0.457
6.683
Token Embeddings
0.084
0.004
0.044
0.038
−0.007
0.015
0.062
0.646
0.354
0.058
0.679
0.369
6.730
Appendix
Table 32 : Evaluating steering guitar type concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( acoustic guitar ) and negative ( electric guitar ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( jazz genre )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.363
0.049
0.206
0.176
0.010
0.093
0.032
0.140
0.086
0.042
0.231
0.137
6.822
PCI (loc.)
0.297
0.038
0.168
0.146
0.012
0.079
0.034
0.114
0.074
0.040
0.160
0.100
6.801
Text Embeddings
0.014
0.028
0.021
−0.025
0.067
0.021
0.352
0.092
0.222
1.184
0.065
0.624
6.633
Text Embeddings (loc.)
0.058
0.031
0.045
0.005
0.042
0.024
0.107
0.077
0.092
0.245
0.058
0.151
6.760
Token Embeddings
−0.011
0.031
0.010
−0.006
0.043
0.019
0.965
0.170
0.568
0.697
0.076
0.386
6.706
Appendix
Table 33 : Evaluating steering jazz genre concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( jazz song ) and negative ( rock song ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Method ( classical genre )
AUC MuQ ( ↑ )
AUC CLAP ( ↑ )
Smoothness MuQ ( ↓ )
Smoothness CLAP ( ↓ )
Avg Quality ( ↑ )
+
-
avg
+
-
avg
+
-
avg
+
-
avg
PCI
0.239
0.049
0.144
0.117
0.028
0.072
0.036
0.108
0.072
0.038
0.088
0.063
6.813
PCI (loc.)
0.206
0.019
0.113
0.102
0.013
0.057
0.034
0.133
0.084
0.043
0.104
0.073
6.798
Text Embeddings
0.075
−0.006
0.034
−0.009
0.055
0.023
0.068
0.242
0.155
0.515
0.051
0.283
6.642
Text Embeddings (loc.)
0.067
−0.000
0.033
0.010
0.044
0.027
0.114
0.215
0.164
0.319
0.068
0.194
6.783
Token Embeddings
0.079
0.022
0.050
0.055
0.064
0.060
0.043
0.163
0.103
0.035
0.052
0.044
6.732
Appendix
Table 34 : Evaluating steering classical genre concept with Ace-Step. We describe metrics in Section 5.2 . Symbols + and - denote the quality of steering towards positive ( classical song ) and negative ( electronic song ) directions, avg = (pos + neg) / 2. Best and second results are highlighted.
Figure 16 : Alignment–preservation curves for piano steering. Used for computing AUC values in Tab. 26 .
Figure 17 : Alignment–preservation curves for mood steering. Used for computing AUC values in Tab. 28 .
Figure 18 : Alignment–preservation curves for tempo steering. Used for computing AUC values in Tab. 27 .
Figure 19 : Alignment–preservation curves for vocal gender steering. Used for computing AUC values in Tab. 29 .
Figure 20 : Alignment–preservation curves for vocal style steering. Used for computing AUC values in Tab. 30 .
Figure 21 : Alignment–preservation curves for violin steering. Used for computing AUC values in Tab. 31 .
Figure 22 : Alignment–preservation curves for guitar type steering. Used for computing AUC values in Tab. 32 .
Figure 23 : Alignment–preservation curves for jazz genre steering. Used for computing AUC values in Tab. 33 .
Figure 24 : Alignment–preservation curves for classical genre steering. Used for computing AUC values in Tab. 34 .
Figure 25 : Mel spectrograms for audio concept modulation. Each row presents the result of the steering audio generation process with SAEs. The central column ( α=0 ) is the unsteered baseline.
Method
Steering Two Concepts
Steering Three Concepts
P + V
P + FV
A + FV
P + FV
J + T
P + V + J
A + FV + M
P + V + T
A + FV + M
AUC MuQ
AUSteer all
+0.057
+0.001
+0.022
+0.013
+0.090
+0.062
+0.005
+0.048
+0.045
CAA all
+0.048
+0.011
+0.028
+0.008
+0.068
+0.047
+0.026
+0.029
+0.032
AUSteer loc
+0.099
+0.018
+0.029
+0.039
+0.086
+0.088
+0.021
+0.069
+0.051
CAA loc
+0.094
+0.034
+0.052
+0.027
+0.103
+0.081
+0.035
+0.061
+0.039
Appendix
Table 35: Multi-concept steering performance measured with average Area Under LPAPS-MUQ and LPAPS-CLAP Curves. Best result per combination of concepts is bolded . Single column denotes steering on combinations of concepts, where P = Piano, V = Violin, FV = Female Vocal, A = Acoustic Guitar, J = Jazz Music, M = Brighter Mood, FV = Male Vocal, T = Slower Tempo, and M = Darker Mood.
Controllable music retrieval lets users find music that is, for example, more ambient, less distorted, or without guitar while preserving the other semantic content of an original seed query. Sparse autoencoders (SAEs) are a promising interface for this kind of concept-level control, but a key problem remains: given a free-form text concept, which sparse features should be edited? In shared multimodal embedding spaces, standard attribution methods often select neurons that match the concept's wording but not the audio examples that express it. This leads to weak or unstable edits: relevant features are missed when concepts are distributed across neurons, while others are selected due to text alignment rather than audio-side structure. We address this with a lightweight, training-free method that recovers a sparse set of audio features whose decoded representation reconstructs the target concept while remaining consistent with audio-space geometry. This reframes concept attribution as a sparse inversion problem rather than a text-side neuron-ranking heuristic. The method requires neither paired audio-text supervision nor SAE retraining. We evaluate this approach in steerable music retrieval and show that the recovered supports align more closely with concept-bearing audio examples and achieve a stronger trade-off between edit strength and preservation than alignment baselines, enabling more precise concept amplification and suppression with reduced drift on preservation metrics.
Julien Guinot, Alain Riou, Elio Quinton +1
Music & Audio Machine Learning Lab, Universal Music Group, London, U.K. · Centre for Digital Music, Queen Mary University of London, U.K.
Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution. However, training audio diffusion models remains computationally expensive, and most existing pipelines still rely on static optimization recipes that treat the relative importance of training signals as fixed throughout learning. In this work, we argue that a major source of inefficiency lies in the evolving balance between semantic acquisition and generation-oriented refinement. Early training places stronger emphasis on acquiring condition-aligned semantic structure and coarse global organization, whereas later training increasingly emphasizes temporal consistency, perceptual fidelity, and fine-detail refinement. To characterize this evolving balance, we introduce a progress-based regime variable derived from the training-time slope of an SSL-space discrepancy, which measures semantic progress during training. Based on this signal, we develop three complementary stage-aware mechanisms: decayed SSL guidance for early semantic bootstrapping, self-adaptive timestep sampling driven by the regime variable, and structure-aware regularization activated from convergent grouped organization in parameter space. We evaluate these mechanisms on text-conditioned audio generation and audio-conditioned super-resolution. Across both settings, the proposed stage-aware strategies improve convergence behavior and yield gains on the primary generation and spectral reconstruction metrics over standard static baselines. These results support the view that efficient audio diffusion training can benefit from treating external guidance, internal organization, and optimization emphasis as stage-dependent components rather than fixed ingredients.
Xuanhao Zhang, Chang Li
China Pharmaceutical University · University of Science and Technology of China
Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in achieving fine-grained, interpretable control over discrete signal attributes. This paper investigates the mechanistic interpretability of the Multitrack Music Transformer (MMT) and proposes a framework for deterministic attribute modulation without retraining to bridge this gap via inference-time activation steering. Utilizing the Difference-in-Means (DiffMean) methodology, we isolate latent directions for signal attributes, specifically Pitch and Duration, within the residual stream. We validate the Linear Representation Hypothesis in this domain, achieving high correlation between steering magnitude and attribute shift. To address the inherent feature entanglement in multi-attribute steering, we introduce a Dual Steering framework utilizing Gram-Schmidt Orthogonalization. Experimental results demonstrate that this geometric decoupling reduces conceptual interference and signal degradation compared to naive vector addition, enabling independent deterministic control even against strong autoregressive conditioning.
Ioannis Prokopiou, Pantelis Vikatos, Maximos Kaliakatsos-Papakostas +2
Athens University of Economics and Business, Athens, Greece · Orfium Research, Athens, Greece · Hellenic Mediterranean University, Chania, Greece +2