Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-passed counterpart, is a long-standing problem in audio processing. In this work, we extend recent advances in neural architectures by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.
Figures & tables
Figure 1: Bandwidth extension as token completion. The codec is frozen, so it ingests the same 48 kHz representation whatever the input bandwidth b . The transformer completes band-limited codes into full-band ones autoregressively over quantizer depth. Only the first predicted level c~1 is decoded, and its content is spliced above the cutoff onto the ground-truth low band, which reaches the output unchanged.
MUSDB18 (in domain)
OrchideaSOL (out of domain)
Input
System
ViSQOL ↑
Mel ↓
STFT ↓
Wav. ↓
SI-SDR ↑
ViSQOL ↑
Mel ↓
STFT ↓
Wav. ↓
SI-SDR ↑
8 kHz
A2SB
2.56
1.14
3.77
2.17
14.89
2.66
1.06
3.23
1.07
24.58
UniverSR
2.77
0.93
2.73
2.25
14.23
2.91
0.83
2.49
1.48
24.91
Ours (SpectroStream)
2.99
0.77
2.49
2.49
14.01
2.71
0.98
2.68
1.16
26.73
Ours (DAC)
2.88
0.76
2.46
2.28
14.66
2.89
0.90
2.56
1.20
26.51
16 kHz
A2SB
3.06
0.63
2.38
1.17
20.73
2.77
0.63
2.37
0.29
36.26
Table 1: Objective metrics at every input rate, with the two corpora grouped: MUSDB18 (in domain) left, OrchideaSOL (out of domain) right. Both of our systems are the single multi-rate predictor, on the codec named. Bold marks the best of the four systems on the three metrics that can arbitrate here. Waveform ℓ1 ( 10−2 ) and SI-SDR (dB) are given for completeness.
Condition
Median
[Q1,Q3]
Hidden reference
100
[100,100]
Band-limited anchor
32
[18.75,49]
A2SB [ 17 ]
70
[59.75,79.25]
Ours (DAC)
67
[53.75,76.25]
Table 2: MUSHRA scores (0–100) at 8 kHz input on MUSDB18, from 10 listeners over 12 excerpts, given as median and interquartile range.
Input
Anchor
Per-rate
Multi-rate
+ rate emb.
Ceiling
8 kHz
1.59
3.01
2.99
2.98
4.55
16 kHz
1.92
3.43
3.40
3.39
4.48
24 kHz
2.98
3.93
3.94
3.93
4.51
32 kHz
4.01
4.35
4.33
4.30
4.56
Table 3: ViSQOL on MUSDB18 across input rates, SpectroStream predictors. Per-rate: four predictors, one per rate. Multi-rate: the single predictor we propose. + rate emb.: the same, with a learned embedding of the cutoff. Anchor: the band-limited input scored as it stands, with nothing synthesized. Ceiling: the true full-band codes decoded at full depth, i.e. perfect prediction at the same bit rate.
School of Computing, National University of Singapore · Department of Statistics & Data Science, National University of Singapore · College of Computing and Data Science, Nanyang Technological University