cs.SDSep 28, 2026

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

Authors: David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Organizations: Princeton University · The Ohio State University · Symbal AI

Abstract

Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.

Figures & tables

Explore similar work

CardsList
  1. PHALAR: Phasors for Learned Musical Audio Representations

    May 5, 2026Davide Marincione, Michele Mancusi, Giorgio Strano +4Audio Representation LearningRepresentation Learning

  2. STEMMA: Song-to-Stem Multi-Audio Reasoning for Large Audio Language Models

    Oct 8, 2026Hoyeol Sohn, Wonil Kim, Keunhyoung Kim +7Audio ReasoningAudio-Language Models

  3. Steering dense music retrieval with open-vocabulary concept discovery

    Aug 9, 2026Julien Guinot, Alain Riou, Elio Quinton +1Feature AttributionSparse Autoencoders