cs.SDSep 18, 2026

Training Music Sample Identification Models on Real Sample Pairs

Authors: R. Oguz Araz, Joan Serrà, Xavier Lizarraga-Seijas, Xavier Serra, Yuki Mitsufuji, Dmitry Bogdanov

Organizations: Music Technology Group, Universitat Pompeu Fabra, Spain · Sony Europe · Sony CTC America

Abstract

Sample identification (SI) is the task of matching pairs of tracks, where one track is created by musically transforming an element of the other. In the absence of sample annotations at scale, the dominant training paradigm has depended on artificially creating sample pairs. Although a recently released dataset provides annotations of real sample pairs at scale, an effective training recipe is missing. In this work, we present SI Embeddings (SIE), an SI model that achieves state-of-the-art results on three benchmarks, including a large-scale test set. We show that the previous state of the art trained on artificial pairs generalizes only partially to real pairs, and that its training data limits its performance. We also show that real pairs do not fully account for SIE's performance: its architecture and training recipe contribute substantially. We provide the first fully supervised training recipe for real-world SI, establishing a strong foundation for future research in the field.

Figures & tables

Explore similar work

Sep 14, 2026cs.SD

Building a Dataset for Music Sample Identification

Sample identification (SI) is the task of matching an element of a musical work to its musically transformed versions used to create new works. The task has received little attention and lacks large-scale publicly available data. In this work, we mine sampling annotations from a music database and split them for training and evaluation. The resulting dataset is nearly three orders of magnitude larger than the existing SI benchmarks, with training, validation, and test sets of 114 k, 6 k, and 10 k tracks. We find that naively splitting the annotations places the same tracks in different sets. To avoid this, we construct a graph from the annotations and split it over connected components. We further find that a single mega-component contains half of the annotations, making component-wise splitting incompatible with balanced splits; we trim it, yielding a leakage-aware pipeline. We share the dataset for non-commercial scientific research purposes only and make the data-analysis and splitting code publicly available. We hope that our work fosters research on SI.
May 22, 2026cs.SD

Music Transcription with (Almost) No Supervision

Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions. Meanwhile, vast quantities of unpaired audio recordings and symbolic scores are freely available but have gone unused. We adopt a cycle-consistent translation framework in which a small amount of paired data acts as a minimal anchor, unlocking the full potential of the unpaired pool. We find that: unpaired data yields surprisingly large gains, especially under limited supervision; unpaired audio contributes more than unpaired scores; incorporating unlabeled audio from a new instrument during training improves transcription for that instrument without any paired supervision. Together, these results suggest that scaling unpaired data offers a practical path toward high-quality transcription for instruments where labeled data remains scarce.
Jun 15, 2026cs.SD

Turning music identification into a neural forward pass

Search, a foundational operation in computer science, maps a query to a matching item in a collection. It is typically implemented as a System-2 like, rule-based pipeline in which a key is computed, an index is probed, and candidates are verified. By contrast, human recognition resembles a System-1 like, associative model of identity recovery, in which even partial cues can trigger a recall without explicitly enumerating, ranking, or even accessing discrete candidates. Here, we show that music sound identification, a difficult search problem, can be performed in a single neural feed-forward pass by a generative transformer. Trained on an audio dataset, the model predicts the corresponding track identifier from a short audio excerpt. This approach surpasses state-of-the-art acoustic fingerprinting, with the largest gains for short audio segments (1 second), demonstrating the method is not only viable but advantageous. Moreover, it reduces external storage to 0.33% of the baseline footprint and improves inference latency by 2.3x (p95). Furthermore, the model can reject queries for unseen tracks, supporting open-set operation while reducing misattribution risk. Using music track identification as an example, this work reframes search, bringing it closer in spirit to human associative recognition and away from algorithmic database lookup.