On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership
Authors: Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos, James C. Davis, George K. Thiruvathukal, Kristen Yeon-Ji Yun, Hao-Wen Dong, Yung-Hsiang Lu
Organizations: Purdue University · Loyola University Chicago · University of Michigan
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39dB, compared with 2.49dB for our strongest baseline. See the demo page at https://benschou.com/notesep.
Figures & tables
Figure 1: Architecture of NoteGrab: NoteGrab extracts one queried note using gated harmonic and ungated percussive components fed into a two-stream version of TFC–TDF U-Net. At three middle stages, a bidirectional cross-attention pair writes information into the opposite stream.
Figure 2: ASO gates the raw magnitudes, predicts relative gains from cross-note context, and normalizes the adjusted weights before reallocating mixture energy. The input channels are described in Sec. 2.4 .
System
SI-SDR
Off-E
Gain
Oct-Conf ↓
Melodyne 5.4.2 †
-3.88 (-4.86)
0.62
0.34
–
Score-Informed NMF
2.49 (2.27)
0.25
0.82
–
Independent Ungated
3.94 (3.59)
0.09
0.62
8.15%
Independent Selective
4.46 (4.24)
0.09
0.62
7.79%
NoteGrab Ungated
4.54 (4.28)
0.11
0.71
8.29%
NoteGrab Selective
5.09 (4.87)
0.12
0.72
5.19%
Table 1: SCNS-Eval results. SI-SDR is per-note median (mean) in dB; Off-E and gain are medians. Higher SI-SDR is better; lower Off-E and Oct-Conf are better; gain should approach one. † Melodyne uses matched notes only.
Figure 3: Octave-overlapping A ♭ 4 and A ♭ 5 notes in an SCNS-Eval piano chord. From top: isolated target, independent NoteGrab estimate, and joint NoteSep output. Dotted lines mark shared partials at 0.83 and 1.66 kHz; h labels denote harmonic number.
System
Pre-edit
Target 2f
Unedited 2f
Final 2f
Score-Informed NMF
2.49
25.73
41.77
43.40
NoteGrab
5.09
31.70
32.87
32.38
NoteSep
7.39
40.09
71.23
71.08
Table 2: SCNS-Eval pre-edit SI-SDR (as in Table 1 ) and 13,892 subsequent edits. Values are medians. Higher is better.
System
SI-SDR ↑
SI-SDRi ↑
SDR ↑
Med.
Mean
Med.
Mean
Med.
Mean
PHENICX-Anechoic, 4 pieces
X-UMX
−4.78
−7.69
4.25
3.33
0.76
–
Score-Inf. X-UMX
−4.79
−6.52
5.12
4.50
1.04
–
NoteGrab
−4.04
−4.69
6.09
6.33
−1.32
–
NoteSep
−1.64
−2.04
8.45
8.98
1.53
–
Table 3: Instrument separation in dB: median and mean.
Data Science Degree Program, National Taiwan University and Academia Sinica, Taipei, Taiwan · Institute of Information Science, Academia Sinica, Taiwan · Rhythm Culture Corporation +1