cs.SDSep 30, 2026

Ghost in the Encoder: Decodable Artist Identity Representations in Lyrics-to-Song Generation

Authors: Arhan Vohra, Choenden Kyirong, Laura Ibáñez-Martínez, Martín Rocamora

Organizations: Music Technology Group, Universitat Pompeu Fabra

Abstract

Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these phenomena have been documented behaviorally on small datasets, little is known about the internal representations that may give rise to them. Prior interpretability work on generative audio has focused on locating semantic concepts such as genre or time signature within model activations. In this work, we show that a trained model can be probed for linearly decodable representations of artist identity from song lyrics alone, without any additional identifiers. Through a controlled case study of ACE-Step 1.5 spanning 2,000 songs across 100 artists, we demonstrate that the artist associated with a given set of lyrics can be identified within the model's internal activations, and that this conditioning signal propagates from the lyric encoder to the diffusion backbone during inference. These findings indicate that lyrics constitute an artist-level conditioning channel not addressed by prompt-side replication safeguards. More broadly, our work highlights how latent-space analysis can be used to audit what generative music models have implicitly learned from their training data.

Figures & tables

Explore similar work

CardsList
  1. SongCraft: Unified Song Generation and Editing with Reconstructive Learning

    Sep 14, 2026Haohe Liu, Varun Nagaraja, Gael Le Lan +5Song Generation

  2. Auditing Training Data in Generative Music Models via Black-Box Membership Inference

    May 28, 2026Yi Chen Liu, Jiawei Yu, Kexin Cao +3Text-To-MusicMembership Inference

  3. HeartMuLa: A Family of Open Sourced Music Foundation Models

    Jan 15, 2026Dongchao Yang, Yuxin Xie, Yuguo Yin +26Song Generation