cs.SDAug 4, 2026

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

Authors: Yizhong GengWenxin FuKecan MaoQifei LiYingming GaoRuimin WangChunfeng WangHao Li+2 more

Organizations: Beijing University of Posts and Telecommunications, Beijing, China · Li Auto, Beijing, China

Abstract

Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.

Explore similar work

CardsList