Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.
Figures & tables
# Encoder Params
SUPERB: linguistic
SUPERB: paralinguistic / speaker
Method
PT data
PR ↓
ASR ↓
KWS ↑
IC ↑
ER ↑
SID ↑
SV ↓
SD ↓
HuBERT Base [ 9 ]
LS-960
95M
6.74
6.97
96.23
96.36
62.44
61.45
6.18
7.19
WavLM Base [ 2 ]
LS-960
95M
5.59
6.42
96.36
96.57
61.35
47.04
8.01
6.41
SPEAR † [ 23 ]
LS-960
95M
4.09
3.90
96.85
98.37
65.78
55.47
6.80
3.23
WavCube [ 20 ]
LS-960
350M
9.91
9.36
97.42
90.41
63.47
42.36
5.86
8.14
SPEAR-Gen Base
LS-960
95M
4.53
3.71
97.47
98.73
69.48
70.23
5.08
3.26
Table 2: Final-layer results on SUPERB. Upper block: understanding-oriented models; lower block: unified understanding–generation models. PT data denotes representation pre-training data. Note: SPEAR-Gen uses external supervision only for TAFA target construction. SPEAR † : reproduced version of SPEAR on LS-960. Best in bold and second-best underlined.
Method
PESQ ↑
ViSQOL ↑
UTMOS ↑
STOI ↑
WER ↓
SIM ↑
WavLM Base [ 2 ]
1.16
1.68
3.37
0.79
2.50
0.473
SPEAR † [ 23 ]
1.20
1.70
3.38
0.81
2.61
0.467
WavCube [ 20 ]
2.93
3.92
3.72
0.95
2.26
0.900
SPEAR-Gen Base
2.94
4.12
3.69
0.95
2.22
0.892
SPEAR-Gen Large
2.98
4.19
3.78
0.96
2.19
0.910
Table 3: Speech resynthesis results on LibriSpeech test-clean.
SUPERB
Resynthesis
MTP target
Gen.
ASR ↓
ER ↑
SV ↓
ViSQOL ↑
WER ↓
SIM ↑
Single-layer
×
3.90
65.78
6.80
1.70
2.61
0.467
TAFA
×
3.83
68.32
6.15
1.86
2.54
0.555
None
✓
20.72
53.83
14.42
4.29
2.23
0.909
Single-layer
✓
3.94
66.70
6.42
4.07
2.24
0.866
LayerAvg
✓
4.35
66.91
6.02
4.06
2.24
0.873
Table 4: Ablation of MTP target and generative modelling (Gen.). MTP target specifies the representation from which MVQ prediction targets are constructed. None denotes training without MTP. LayerAvg denotes uniform layer averaging.
#Decoder Params
SUPERB
Resynthesis
Gen. Obj.
ASR ↓
ER ↑
SV ↓
PESQ ↑
UTMOS ↑
WER ↓
SIM ↑
Mel ℓ1
15M
3.79
68.42
5.41
2.76
3.55
2.25
0.883
Mel ℓ1
30M
3.70
69.28
5.46
2.73
3.61
2.22
0.883
Full-mel FM
30M
3.72
68.92
5.59
2.61
3.58
2.26
0.875
SPEAR-Gen
30M
3.71
69.48
5.08
2.94
3.69
2.22
0.892
Table 5: Ablation of the generative objectives. #Params denotes the total number of parameters in decoders.