Sample identification (SI) is the task of matching pairs of tracks, where one track is created by musically transforming an element of the other. In the absence of sample annotations at scale, the dominant training paradigm has depended on artificially creating sample pairs. Although a recently released dataset provides annotations of real sample pairs at scale, an effective training recipe is missing. In this work, we present SI Embeddings (SIE), an SI model that achieves state-of-the-art results on three benchmarks, including a large-scale test set. We show that the previous state of the art trained on artificial pairs generalizes only partially to real pairs, and that its training data limits its performance. We also show that real pairs do not fully account for SIE's performance: its architecture and training recipe contribute substantially. We provide the first fully supervised training recipe for real-world SI, establishing a strong foundation for future research in the field.
Figures & tables
System
Training Pairs
Dim.
Sample100
SamplePairs
mAP ( ↑ )
mNAR ( ↓ )
mAP ( ↑ )
mNAR ( ↓ )
Van Balen et al. [ 16 ]
None (rule-based)
–
0.390 †
n/a
n/a
n/a
Cheston et al. [ 9 ]
Artificial
2048
0.441 †
n/a
n/a
n/a
SampleID [ 13 ]
Artificial
2048
0.603
7.6
0.450
4.7
SampleID-R (ours)
Real
2048
0.651
6.3
0.511
4.1
SIE (ours)
Real
1024
0.758
3.9
0.701
2.4
Table 1: Comparison of SI systems on the Sample100 and SamplePairs benchmarks. Dim. denotes embedding dimensionality. SampleID-R is our retraining of SampleID on real pairs. † denotes scores cited from corresponding publications; n/a denotes scores that are neither published nor computable by us due to lack of available implementation.
Model
mAP ( ↑ )
mNAR ( ↓ )
SampleID [ 13 ]
0.226
± 0.007
19.4
± 0.5
SampleID-R
0.299
± 0.008
14.3
± 0.4
SIE
0.389
± 0.009
13.0
± 0.4
NMFP [ 1 ]
0.101
± 0.005
35.4
± 0.6
CLEWS [ 15 ]
0.181
± 0.007
32.4
± 0.7
Fish [ 3 ]
0.207
± 0.007
29.9
± 0.6
Table 2: Comparison of models on the WhoSampled130K test set. SampleID-R is our retraining of SampleID on real pairs. The last three models were not trained for sample identification (see text). Numbers following ± are 95% confidence interval half-widths for the mean over 10,444 queries.
Figure 1: Effect of training SIE on nested random subsets of the WhoSampled130K training set pairs, with everything else held fixed. Evaluated on the WhoSampled130K test set. Dashed horizontal lines mark the scores of SampleID, taken from Table 2 .
SM
AD
mAP ( ↑ )
mNAR ( ↓ )
✗
✗
0.320
± 0.008
16.2
± 0.5
✗
✓
0.334
± 0.008
16.1
± 0.5
✓
✗
0.365
± 0.009
13.0
± 0.4
✓
✓
0.389
± 0.009
13.0
± 0.4
Table 3: Ablation of signal manipulation (SM) and audio degradation (AD) during SIE training, evaluated on the WhoSampled130K test set.
Loss
mAP ( ↑ )
mNAR ( ↓ )
Triplet
0.389
± 0.009
13.0
± 0.4
NT-Xent ( τ0=0.01 )
0.332
± 0.008
12.6
± 0.4
NT-Xent ( τ0=0.05 )
0.331
± 0.008
13.0
± 0.4
NT-Xent ( τ0=0.10 )
0.320
± 0.008
13.0
± 0.4
Table 4: Comparison of triplet and NT-Xent losses for training SIE, evaluated on the WhoSampled130K test set. τ0 is the initialization of the learnable temperature.
Segment Duration (s)
mAP ( ↑ )
mNAR ( ↓ )
5
0.384
± 0.009
12.7
± 0.4
6
0.389
± 0.009
13.0
± 0.4
8
0.381
± 0.009
13.1
± 0.4
Table 5: Effect of segment duration on SIE, evaluated on the WhoSampled130K test set. Each model uses its training segment duration at inference, with hop fixed at 3 s and batch size fixed across models.