Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these phenomena have been documented behaviorally on small datasets, little is known about the internal representations that may give rise to them. Prior interpretability work on generative audio has focused on locating semantic concepts such as genre or time signature within model activations. In this work, we show that a trained model can be probed for linearly decodable representations of artist identity from song lyrics alone, without any additional identifiers. Through a controlled case study of ACE-Step 1.5 spanning 2,000 songs across 100 artists, we demonstrate that the artist associated with a given set of lyrics can be identified within the model's internal activations, and that this conditioning signal propagates from the lyric encoder to the diffusion backbone during inference. These findings indicate that lyrics constitute an artist-level conditioning channel not addressed by prompt-side replication safeguards. More broadly, our work highlights how latent-space analysis can be used to audit what generative music models have implicitly learned from their training data.
Figures & tables
Conditioning input
no-CoT
with-CoT
Lyric encoder
lyrics
lyrics
Caption
null
generated from lyrics
BPM / key
null
generated from lyrics
Timbre
null
null
Table 1 : Conditioning inputs to the DiT under the two inference configurations. Both pass the same lyrics to the lyric encoder; with-CoT additionally activates the chain-of-thought module, which derives a caption and infers BPM/key. All other inputs are held fixed.
Figure 1 : Artist probe accuracy rises through the lyric encoder layers. Error bars show the 95% confidence interval.
Figure 2 : Genre-wise mean artist decodability from the lyric encoder’s output.
Figure 3 : Cross-attention injects decodable artist representations into the first few layers of the DiT audio-token hidden states; these become progressively less decodable in deeper layers.
MuQ-MuLan
CLAP
Originals
no-CoT
with-CoT
Originals
no-CoT
with-CoT
Full dataset (100 artists, chance = 0.01)
All genres
0.449
0.109
0.101
0.224
0.043
0.042
Within-genre (20 artists, chance = 0.05)
Hip-hop
0.445
0.180
0.183
0.273
0.115
0.130
Rock
0.659
0.208
0.130
0.394
0.108
0.073
Table 2 : Artist probe accuracy on generated audio embeddings. All results p<0.001 .
K
with-CoT
no-CoT
Chance
10
0.055
0.022
0.005
25
0.088
0.045
0.013
50
0.137
0.073
0.025
100
0.210
0.131
0.050
200
0.304
0.204
0.100
400
0.449
0.333
0.200
Table 3 : CoverID Retrieval @K via CLEWS. Random baseline = K/N where N=2,000 .