cs.SDSep 3, 2026

Neural Music Enhancement with Dual Time-Frequency Spectral Representations for Prediction and Discrimination

Authors: Fei LiuYang AiZhen-Hua Ling

Organizations: National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China

Abstract

Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT's fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT's log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.

Explore similar work

CardsList
  1. Low-Latency Neural Models for Real-Time Music Enhancement

    Jul 14, 2026Emmanouil Karystinaios, Jonathan Greif, David Nadrchal +2Speech EnhancementMusic Editing

  2. Latent Fourier Transform

    Apr 20, 2026Mason Wang, Cheng-Zhi Anna HuangMusic GenerationFourier