Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-passed counterpart, is a long-standing problem in audio processing. In this work, we extend recent advances in neural architectures by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.
Figures & tables
Figure 1: Bandwidth extension as token completion. The codec is frozen, so it ingests the same 48 kHz representation whatever the input bandwidth b . The transformer completes band-limited codes into full-band ones autoregressively over quantizer depth. Only the first predicted level c~1 is decoded, and its content is spliced above the cutoff onto the ground-truth low band, which reaches the output unchanged.
MUSDB18 (in domain)
OrchideaSOL (out of domain)
Input
System
ViSQOL ↑
Mel ↓
STFT ↓
Wav. ↓
SI-SDR ↑
ViSQOL ↑
Mel ↓
STFT ↓
Wav. ↓
SI-SDR ↑
8 kHz
A2SB
2.56
1.14
3.77
2.17
14.89
2.66
1.06
3.23
1.07
24.58
UniverSR
2.77
0.93
2.73
2.25
14.23
2.91
0.83
2.49
1.48
24.91
Ours (SpectroStream)
2.99
0.77
2.49
2.49
14.01
2.71
0.98
2.68
1.16
26.73
Ours (DAC)
2.88
0.76
2.46
2.28
14.66
2.89
0.90
2.56
1.20
26.51
16 kHz
A2SB
3.06
0.63
2.38
1.17
20.73
2.77
0.63
2.37
0.29
36.26
Table 1: Objective metrics at every input rate, with the two corpora grouped: MUSDB18 (in domain) left, OrchideaSOL (out of domain) right. Both of our systems are the single multi-rate predictor, on the codec named. Bold marks the best of the four systems on the three metrics that can arbitrate here. Waveform ℓ1 ( 10−2 ) and SI-SDR (dB) are given for completeness.
Condition
Median
[Q1,Q3]
Hidden reference
100
[100,100]
Band-limited anchor
32
[18.75,49]
A2SB [ 17 ]
70
[59.75,79.25]
Ours (DAC)
67
[53.75,76.25]
Table 2: MUSHRA scores (0–100) at 8 kHz input on MUSDB18, from 10 listeners over 12 excerpts, given as median and interquartile range.
Input
Anchor
Per-rate
Multi-rate
+ rate emb.
Ceiling
8 kHz
1.59
3.01
2.99
2.98
4.55
16 kHz
1.92
3.43
3.40
3.39
4.48
24 kHz
2.98
3.93
3.94
3.93
4.51
32 kHz
4.01
4.35
4.33
4.30
4.56
Table 3: ViSQOL on MUSDB18 across input rates, SpectroStream predictors. Per-rate: four predictors, one per rate. Multi-rate: the single predictor we propose. + rate emb.: the same, with a learned embedding of the cutoff. Anchor: the band-limited input scored as it stands, with nothing synthesized. Ceiling: the true full-band codes decoded at full depth, i.e. perfect prediction at the same bit rate.
Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-passed counterpart, is a long-standing problem in audio processing. In this work, we extend recent advances in neural architectures by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.
Benoît Giniès, Xiaoyu Bie, Olivier Fercoq +1
LTCI, Télécom Paris, Institut Polytechnique de Paris Palaiseau, France
Recent audio restoration increasingly relies on large-scale conditional latent generative modeling, including diffusion, Schrodinger Bridges, and Flow Matching variants, to invert degradations such as bandwidth limitation or noise. We present an analysis of the performance of various state-of-the-art methods compared to simple arithmetic transformations in the latent spaces of multiple neural codecs for musical bandwidth extension. We show that estimating a single transport vector between the clean and degraded latent centroids on a reference set, and adding it to degraded latents, can yield restoration performance competitive with large diffusion models. This suggests, first, that some neural codec latent spaces exhibit structure aligned with audio bandwidth; and second, that in such cases complex conditional models may offer only limited gains over a simple vector addition. We argue that these findings reveal an interesting avenue for future research whereby models could take advantage of the latent space structure in order to offer greater training and parameter efficiency, and overall better performance. Additionally, we propose to consider this simple arithmetic transformation as a baseline for music bandwidth extension research, as it allows an assessment of the contribution of learnable parameters towards restoration performance.
Hendrik Vincent Koops, Hao Hao Tan, Elio Quinton
Music & Audio Machine Learning Lab · Universal Music Group, London, U.K.
Bandwidth extension (BWE) aims to recover missing high-frequency content from band-limited speech. Existing methods often formulate BWE as a fixed or predefined bandwidth conversion problem, potentially requiring cutoff-specific models or retraining when the input bandwidth changes. This assumption limits their applicability to practical scenarios where speech may arrive with diverse cutoff frequencies. We propose AnyBand, a unified BWE framework that recasts bandwidth extension as in-context spectral infilling. Motivated by prompt-based zero-shot speech generation, AnyBand conditions high-frequency generation on the observed low-frequency spectrum, using the available band as a frequency-domain prompt that conveys content, speaker, prosodic, and spectral-envelope cues. This formulation enables a single model to perform cutoff-conditioned generation over a continuous range of input bandwidths. AnyBand is trained with missing-band conditional flow matching and an Easy-to-Balanced cutoff curriculum over continuously sampled cutoff frequencies. To better exploit the spectral prompt, we introduce a frequency-aware Diffusion Transformer that models cross-frequency interactions and long-range temporal dependencies, followed by a physically motivated multi-view adversarial refinement stage to enhance spectral realism, envelope coherence, and harmonic consistency. Experiments on multiple datasets and bandwidth settings show that AnyBand consistently improves spectral reconstruction over existing baselines while achieving competitive perceptual quality across both standard and irregular input cutoffs. Audio samples are available.
Junchuan Zhao, Minh Duc Vu, Bowen Zhang +1
School of Computing, National University of Singapore · Department of Statistics & Data Science, National University of Singapore · College of Computing and Data Science, Nanyang Technological University