cs.SDSep 24, 2026

On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership

Authors: Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos, James C. Davis, George K. Thiruvathukal, Kristen Yeon-Ji Yun, Hao-Wen Dong, Yung-Hsiang Lu

Organizations: Purdue University · Loyola University Chicago · University of Michigan

Abstract

Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39dB, compared with 2.49dB for our strongest baseline. See the demo page at https://benschou.com/notesep.

Figures & tables

Explore similar work

CardsList
  1. SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings

    Sep 27, 2026Yiheng Lu, Hao-Wen DongSpeech SeparationInstrument

  2. Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion

    Aug 2, 2026Wei-Han Hsu, Chih-Cheng Chang, Bo-Yu Chen +2Multi-Instrument Music TranscriptionSpeech Separation