Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.
Figures & tables
R@1
R@5
System
All
Family
All
Family
Trained retrieval
Stembed
56.8±0.9
58.9
78.1
81.2
CIR–MuQ
23.6±0.5
52.4
60.7
81.7
SMTI–MuQ
21.4±0.6
23.7
39.8
44.0
Frozen encoders
Table 1: Three-stem MoisesDB retrieval (percent). Trained rows report three-seed means and sample standard deviation (SD) for All R@1. Other SDs are at most 1.9 points. Frozen rows use one evaluation. Bold marks column maxima, excluding GT-stem references and ablations. “All” searches 2068 stems; “Family” searches only the target’s family.
R@1
R@5
System
All
Family
All
Family
Stembed
38.1±1.1
43.2±1.1
59.0±1.5
68.0±1.2
Predicted count
37.5±0.1
41.4±0.4
56.7±0.7
62.9±0.4
CIR–MuQ
20.2±0.3
42.6±0.5
45.1±0.3
70.0±0.7
SMTI–MuQ
15.8±0.7
19.3±0.5
29.4±0.5
36.8±0.6
Table 2: Three-stem retrieval on Slakh2100 using models selected on MoisesDB. Values are mean ± sample SD (%) across three training seeds. The gallery contains 1458 stems, with search protocols defined in Table 1 . Bold marks the best mean in each column.