cs.SDSep 24, 2026

On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership

Authors: Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos, James C. Davis, George K. Thiruvathukal, Kristen Yeon-Ji Yun, Hao-Wen Dong, Yung-Hsiang Lu

Organizations: Purdue University · Loyola University Chicago · University of Michigan

Abstract

Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39dB, compared with 2.49dB for our strongest baseline. See the demo page at https://benschou.com/notesep.

Figures & tables

Explore similar work

Sep 27, 2026cs.SD

SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings

Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask prediction. We introduce SCISSOR (Score-Conditioned Instrument Source Separation for Orchestral Recordings), which uses the score to form a frame-wise query for each source. Each query matches a shared audio representation, and a softmax over instrument and background slots jointly assigns overlapping time-frequency evidence. The queries retain instrument identity even when notes are missing from the score. After training on SynthSOD and a small set of URMP and PHENICX-Anechoic recordings, SCISSOR achieves the highest average SDR on held-out real recordings. With SynthSOD-only training, it leads on SynthSOD and zero-shot PHENICX-Anechoic, and improves on its audio-only control on zero-shot URMP. SCISSOR also degrades less under score corruption than the evaluated score-based baselines.
Apr 22, 2026cs.SD

From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR

We propose a new approach for a practical two-stage Optical Music Recognition (OMR) pipeline, with a particular focus on its second stage. Given symbol and event candidates from the visual pipeline, we decode them into an editable, verifiable, and exportable score structure. We focus on complex polyphonic staff notation, especially piano scores, where voice separation and intra-measure timing are the main bottlenecks. Our approach formulates second-stage decoding as a structure decoding problem and uses topology recognition with probability-guided search (BeadSolver) as its core method. We also describe a data strategy that combines procedural generation with recognition-feedback annotations. The result is a practical decoding component for real OMR systems and a path to accumulate structured score data for future end-to-end, multimodal, and RL-style methods.
Aug 2, 2026cs.SD

Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion

Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alternative separate-and-detect formulation, where a drum source separation front end first produces five editable drum stems, and a fixed onset detector then converts each stem into symbolic events. The separator is built on a five-stem latent diffusion model that jointly generates kick, snare, toms, hi-hats, and cymbals in a compact VAE latent space. We further study two training-only auxiliary branches--an onset branch (OB) and a timbre branch (TB)--which shape the separator during learning but are discarded at inference. Trained on synthetic drum multitracks and evaluated on MDB Drums and ENST-Drums, the proposed pipeline consistently improves over a strong U-Net-based drum separation baseline in overall transcription F1. It also outperforms a representative end-to-end ADT system on kick and snare F1 under our evaluation protocol, while additionally providing separated audio stems. The ablation results show that OB gives the most stable transcription gains, whereas TB changes the trade-off between reconstruction, acoustic stem quality, and onset detection. These results suggest that generative drum demixing can serve not only as a source separation model, but also as a practical front end for interpretable drum transcription.