cs.SDSep 27, 2026

SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings

Authors: Yiheng Lu, Hao-Wen Dong

Organizations: The Chinese University of Hong Kong, Shenzhen · University of Michigan

Abstract

Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask prediction. We introduce SCISSOR (Score-Conditioned Instrument Source Separation for Orchestral Recordings), which uses the score to form a frame-wise query for each source. Each query matches a shared audio representation, and a softmax over instrument and background slots jointly assigns overlapping time-frequency evidence. The queries retain instrument identity even when notes are missing from the score. After training on SynthSOD and a small set of URMP and PHENICX-Anechoic recordings, SCISSOR achieves the highest average SDR on held-out real recordings. With SynthSOD-only training, it leads on SynthSOD and zero-shot PHENICX-Anechoic, and improves on its audio-only control on zero-shot URMP. SCISSOR also degrades less under score corruption than the evaluated score-based baselines.

Figures & tables

Explore similar work

CardsList
  1. On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership

    Sep 24, 2026Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos +5Speech SeparationSeparation

  2. TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data

    Sep 1, 2026Jianhuai Hu, Yashan Wang, Shangda Wu +7Text-To-MusicTranscription