Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-annotated corpus of clean per-stem combo recordings on standards. We introduce JazzSAMBA (Jazz Synchronous and Asynchronous Multi-take Band Audio) to fill this gap: the first originally recorded jazz-combo multitrack dataset of standards with asynchronous (overdubbed) and synchronous (live ensemble) protocols, preferred and alternate takes chosen by the musicians, and timed annotations for bars, chords, sections, and soloists. JazzSAMBA covers 76 standards by eight musicians on drums, bass, piano, trumpet, and saxophone, with per-stem audio, mixtures, and MIDI. It can support chart-conditioned accompaniment, combo source separation, and form-aware music information retrieval. We demonstrate the dataset on two tasks: a jazz combo source-separation baseline and a chart-conditioned accompaniment ablation. The dataset, code, and samples are linked from the project demo page.
Figures & tables
Figure 1: JazzSAMBA overview ( 50 asynchronous / 26 synchronous songs; about 8.1 h preferred-take, about 16.1 h both tiers). Labels 1 and 2 are take tiers. Asynchronous takes overdub in sequence (drums first; dotted arrows); synchronous takes record the whole ensemble together (dashed lines). Preferred mixtures are annotated for bars, chords, sections, and soloists; alternate async annotations are copied, while sync takes are labeled separately.
Corpus
Stems
Takes
Dual
Chords
Sections
Soloists
MIDI
MUSDB18 [ 19 ]
✓
MoisesDB [ 17 ]
✓
Slakh2100 [ 15 ]
✓
✓
Jazz Trio Database [ 5 ]
✓ †
✓ ∗
Jazz Harmony [ 7 ]
✓
✓
Weimar Jazz Database [ 18 ]
✓
✓
✓
✓ ∗
Table 1: JazzSAMBA vs. selected related corpora.
Protocol
Stems
Annotations
GT MIDI
Clean
Bleed
Timed
Shared
Drums
Piano
Asynchronous
✓
✓
✓
✓
✓
Synchronous
✓
✓
✓
Table 2: Asynchronous vs. synchronous protocol contents. Shared marks preferred labels copied to the alternate tier. GT denotes ground-truth; remaining MIDI is from MuScriptor AMT [ 20 ] . For stems with bleed, we also ship a Wiener-style de-bleeded version [ 10 ] . All audio is 48 kHz.
Model
FT
Drums
Bass
Piano
Horns
HT-Demucs 6s
—
13.01
15.07
4.92
—
JazzSAMBA
15.09
17.03
9.42
14.62
ChoraleBricks
—
15.02
—
8.98
Open-Unmix
—
5.24
5.63
—
—
JazzSAMBA
8.37
9.74
6.07
9.38
ChoraleBricks
—
5.63
—
− 0.66
Table 3: Asynchronous-test SI-SDR ↑ (dB), full-track ( n=5 ). Dashes mark unfair zero-shot or ChoraleBricks cells (no matching stem supervision). We bold the best reportable score within each model.
COCOLA ↑
Beat F1 ↑
Protocol
Condition
D
B
P
H
D
B
P
H
Asynchronous
Ground Truth
55.6
62.1
63.4
64.5
1.00
0.80
0.48
0.29
Listen-Only
52.2
60.4
59.8
61.4
0.23
0.16
0.13
0.02
+ Sections
50.7
59.8
59.7
61.9
0.29
0.09
0.13
0.05
+ Chords
52.5
60.6
60.0
62.3
0.30
0.09
0.15
0.06
+ Both
51.4
60.4
59.9
61.7
0.34
0.10
0.12
0.04
Table 4: Accompaniment scores on 60 s free-run from form landmarks (asynchronous test n=5 , synchronous test n=3 ; mean over landmarks). D, B, P, and H denote drums, bass, piano, and horns. Ground Truth scores the annotated target stem against the condition mix. We bold the best score per column among chart conditions (excluding Ground Truth).
Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval. Our results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-k results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification. Project page: https://github.com/cagries/tipofmyear.
Çağrı Eser
Department of Computer Engineering, Middle East Technical University, Ankara, Turkey.
Sample identification (SI) is the task of matching an element of a musical work to its musically transformed versions used to create new works. The task has received little attention and lacks large-scale publicly available data. In this work, we mine sampling annotations from a music database and split them for training and evaluation. The resulting dataset is nearly three orders of magnitude larger than the existing SI benchmarks, with training, validation, and test sets of 114 k, 6 k, and 10 k tracks. We find that naively splitting the annotations places the same tracks in different sets. To avoid this, we construct a graph from the annotations and split it over connected components. We further find that a single mega-component contains half of the annotations, making component-wise splitting incompatible with balanced splits; we trim it, yielding a leakage-aware pipeline. We share the dataset for non-commercial scientific research purposes only and make the data-analysis and splitting code publicly available. We hope that our work fosters research on SI.
R. Oguz Araz, Xavier Lizarraga, Xavier Serra +1
Music Technology Group, Universitat Pompeu Fabra, Spain
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.
Simon Rouard, Michael Krause, Axel Roebel +2
Kyutai · Mirelo AI · UMR STMS, IRCAM-CNRS Sorbonne Univ.