Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-passed counterpart, is a long-standing problem in audio processing. In this work, we extend recent advances in neural architectures by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.
Figures & tables
Fig. 1: HP-codec , our structure-informed disentangled codec. It is divided in two branches operating at different sampling rates: a 16 kHz branch and a 48 kHz branch. Each branch contains parallel RVQs which are composed of a harmonic section and a percussive section.
Fig. 2: HP-codecX , our bandwidth extension model. It connects the 16 kHz representation, extracted from the input, to the 48 kHz decoder, through an audio language model organized into two sub-models: a harmonic estimator and a percussive estimator.
SR=16kHz
Mel ↓
STFT ↓
Waveform ↓ ( 10−2 )
HP-codec
0.79 [0.77,0.78]
2.26 [2.25,2.28]
4.8 [4.8,4.9]
soft-dis-DAC [ 30 ]
0.70 [0.69,0.70]
2.11 [2.10,2.13]
4.1 [4.1,4.2]
DAC [ 15 ]
0.82 [0.81,0.82]
2.32 [2.30,2.34]
5.1 [5.0,5.2]
TABLE I: Reconstruction metrics (and 95% CI) for HP-codec. Soft-dis-DAC is a modified version of [ 30 ] in which the structure-informed sections (Harmonic and Percussive RVQs) are removed. DAC-16kHz and DAC-48kHz denotes DAC models [ 15 ] retrained on our dataset.
Fig. 3: Reconstruction metrics for HP-codec as a function of the spectral composition of the input and the RVQ sections used for reconstruction: harmonic tokens ( H ), percussive tokens ( P ), or both ( H+P ). These plots are to be read as follows: the larger the enclosed area, the better the reconstruction. The Global configuration ( H+P ) consistently performs best, reflecting its larger representational capacity. The Harmonic section achieves the best performance on harmonic inputs at 16 kHz, while the Percussive section performs best on percussive inputs at 48 kHz. This pattern mirrors the natural spectral distribution of audio, with harmonic energy concentrated at low frequencies and percussive energy extending toward higher frequencies.
Fig. 4: Variation in harmonic and percussive energy concentration between reconstructions using only harmonic ( H ) or only percussive ( P ) tokens and the full reconstruction ( H+P ).
Fig. 5: Spectrograms of estimated signal drawn from Apollo ( APO ), AudioSR ( ASR ), A2SB ( A2S ), UniverSR ( USR ), and HP-codecX ( HPX ).
Mel ↓
STFT ↓
Waveform ↓ ( 10−2 )
Apollo [ 47 ]
0.84 [0.83,0.85]
2.44 [2.41,2.47]
4.8 [4.8,4.9]
AudioSR [ 42 ]
1.77 [1.74,1.80]
3.60 [3.55,3.64]
6.9 [6.8,7.0]
A2SB [ 44 ]
0.63 [0.62,0.64]
2.65 [2.61,2.69]
1.2 [1.2,1.3]
UniverSR [ 43 ]
1.32 [1.29,1.34]
3.53 [3.48,3.58]
0.9 [0.9,1.0]
HP-CodecX
0.45 [0.44,0.46]
1.89 [1.87,1.91]
1.2 [1.2,1.3]
TABLE II: Objective bandwidth extension metrics (and 95% CI) for Apollo (44.1 kHz), AudioSR (48 kHz), A2SB (44.1 kHz), UniverSR (48 kHz) and HP-codecX (48 kHz) on MUSDB18 test set.
Fig. 6: Results of the perceptual evaluation (with median, first and third quartiles). The MUSHRA test compared Apollo ( APO ) and A2SB ( A2S ) models to HP-codecX ( HPX ). Reference signals ( Ref ) and anchor signals ( Anc ) were also evaluated.
Fig. 7: Objective reconstruction metrics calculated on estimated signals. These metrics have been computed on Out-of-Domain test datasets: ENST-Drums, Medley-solos-DB, OrchideaSOL, Monophonic, Polyphonic, ESC-50 and VCTK. The Apollo ( APO ) and A2SB ( A2S ) metrics are calculated at 44.1 kHz, while the AudioSR ( ASR ), UniverSR ( USR ) and HP-codecX ( HPX ) metrics have been calculated at 48 kHz.
HP-CodecX
EXP 1
EXP 2
Spec. informed Training
✓
×
×
Multi-transformer
✓
✓
×
Multi-RVQ
✓
✓
✓
Frequency Branches
✓
✓
✓
Mel ↓
0.45 [0.44,0.46]
0.48 [0.47,0.48]
0.48 [0.47,0.48]
STFT ↓
1.89 [1.87,1.91]
1.97 [1.92,1.99]
1.97 [1.95,1.99]
TABLE III: Ablation Study: Objective Bandwidth Extension Metrics (and 95% CI) for HP-codecX and five additional configurations obtained by progressively removing architectural and design choices introduced in this work. EXP 5 corresponds to a configuration in which a simple DAC model is used to predict full-bandwidth tokens directly from lossy tokens.
General Music
Harmonic
Medley-solos-DB
OrchideaSOL
Monophonic
Polyphonic
Mel ↓
H
0.42 [0.41,0.43]
0.40 [0.39,0.41]
0.41 [0.39,0.42]
0.51 [0.50,0.52]
P
0.41 [0.40,0.42]
0.39 [0.38,0.40]
0.44 [0.43,0.46]
0.53 [0.52,0.53]
H+P
0.39 [0.39,0.40]
0.38 [0.37,0.39]
0.42 [0.41,0.44]
0.51 [0.50,0.52]
STFT ↓
TABLE IV: Objective Bandwidth Extension Metrics (and 95% CI) across various datasets, using only harmonic ( H ) tokens, only percussive ( P ) tokens, or the combination of both ( H+P ).
Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-passed counterpart, is a long-standing problem in audio processing. In this work, we extend recent advances in neural architectures by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.
Benoît Ginies, Olivier Fercoq, Gaël Richard
LTCI, Télécom Paris, Institut Polytechnique de Paris, Palaiseau, France
Recent audio restoration increasingly relies on large-scale conditional latent generative modeling, including diffusion, Schrodinger Bridges, and Flow Matching variants, to invert degradations such as bandwidth limitation or noise. We present an analysis of the performance of various state-of-the-art methods compared to simple arithmetic transformations in the latent spaces of multiple neural codecs for musical bandwidth extension. We show that estimating a single transport vector between the clean and degraded latent centroids on a reference set, and adding it to degraded latents, can yield restoration performance competitive with large diffusion models. This suggests, first, that some neural codec latent spaces exhibit structure aligned with audio bandwidth; and second, that in such cases complex conditional models may offer only limited gains over a simple vector addition. We argue that these findings reveal an interesting avenue for future research whereby models could take advantage of the latent space structure in order to offer greater training and parameter efficiency, and overall better performance. Additionally, we propose to consider this simple arithmetic transformation as a baseline for music bandwidth extension research, as it allows an assessment of the contribution of learnable parameters towards restoration performance.
Hendrik Vincent Koops, Hao Hao Tan, Elio Quinton
Music & Audio Machine Learning Lab · Universal Music Group, London, U.K.
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.
Qiaoyu Yang, Lixing He, Binyue Deng +1
Georgia Institute of Technology, Atlanta, United States · The Chinese University of Hong Kong, Hong Kong, China · Tencent Music Entertainment, Shenzhen, China