cs.SDAug 9, 2026

Steering dense music retrieval with open-vocabulary concept discovery

Authors: Julien GuinotAlain RiouElio QuintonGyörgy Fazekas

Organizations: Music & Audio Machine Learning Lab, Universal Music Group, London, U.K. · Centre for Digital Music, Queen Mary University of London, U.K.

Abstract

Controllable music retrieval lets users find music that is, for example, more ambient, less distorted, or without guitar while preserving the other semantic content of an original seed query. Sparse autoencoders (SAEs) are a promising interface for this kind of concept-level control, but a key problem remains: given a free-form text concept, which sparse features should be edited? In shared multimodal embedding spaces, standard attribution methods often select neurons that match the concept's wording but not the audio examples that express it. This leads to weak or unstable edits: relevant features are missed when concepts are distributed across neurons, while others are selected due to text alignment rather than audio-side structure. We address this with a lightweight, training-free method that recovers a sparse set of audio features whose decoded representation reconstructs the target concept while remaining consistent with audio-space geometry. This reframes concept attribution as a sparse inversion problem rather than a text-side neuron-ranking heuristic. The method requires neither paired audio-text supervision nor SAE retraining. We evaluate this approach in steerable music retrieval and show that the recovered supports align more closely with concept-bearing audio examples and achieve a stronger trade-off between edit strength and preservation than alignment baselines, enabling more precise concept amplification and suppression with reduced drift on preservation metrics.

Explore similar work

CardsList
  1. FIGMA: Towards FIne-Grained Music retrievAl

    Jun 4, 2026Nishit Anand, Ashish Seth, Sreyan Ghosh +2Music Understanding