Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.
Figures & tables
# Encoder Params
SUPERB: linguistic
SUPERB: paralinguistic / speaker
Method
PT data
PR ↓
ASR ↓
KWS ↑
IC ↑
ER ↑
SID ↑
SV ↓
SD ↓
HuBERT Base [ 9 ]
LS-960
95M
6.74
6.97
96.23
96.36
62.44
61.45
6.18
7.19
WavLM Base [ 2 ]
LS-960
95M
5.59
6.42
96.36
96.57
61.35
47.04
8.01
6.41
SPEAR † [ 23 ]
LS-960
95M
4.09
3.90
96.85
98.37
65.78
55.47
6.80
3.23
WavCube [ 20 ]
LS-960
350M
9.91
9.36
97.42
90.41
63.47
42.36
5.86
8.14
SPEAR-Gen Base
LS-960
95M
4.53
3.71
97.47
98.73
69.48
70.23
5.08
3.26
Table 2: Final-layer results on SUPERB. Upper block: understanding-oriented models; lower block: unified understanding–generation models. PT data denotes representation pre-training data. Note: SPEAR-Gen uses external supervision only for TAFA target construction. SPEAR † : reproduced version of SPEAR on LS-960. Best in bold and second-best underlined.
Method
PESQ ↑
ViSQOL ↑
UTMOS ↑
STOI ↑
WER ↓
SIM ↑
WavLM Base [ 2 ]
1.16
1.68
3.37
0.79
2.50
0.473
SPEAR † [ 23 ]
1.20
1.70
3.38
0.81
2.61
0.467
WavCube [ 20 ]
2.93
3.92
3.72
0.95
2.26
0.900
SPEAR-Gen Base
2.94
4.12
3.69
0.95
2.22
0.892
SPEAR-Gen Large
2.98
4.19
3.78
0.96
2.19
0.910
Table 3: Speech resynthesis results on LibriSpeech test-clean.
SUPERB
Resynthesis
MTP target
Gen.
ASR ↓
ER ↑
SV ↓
ViSQOL ↑
WER ↓
SIM ↑
Single-layer
×
3.90
65.78
6.80
1.70
2.61
0.467
TAFA
×
3.83
68.32
6.15
1.86
2.54
0.555
None
✓
20.72
53.83
14.42
4.29
2.23
0.909
Single-layer
✓
3.94
66.70
6.42
4.07
2.24
0.866
LayerAvg
✓
4.35
66.91
6.02
4.06
2.24
0.873
Table 4: Ablation of MTP target and generative modelling (Gen.). MTP target specifies the representation from which MVQ prediction targets are constructed. None denotes training without MTP. LayerAvg denotes uniform layer averaging.
#Decoder Params
SUPERB
Resynthesis
Gen. Obj.
ASR ↓
ER ↑
SV ↓
PESQ ↑
UTMOS ↑
WER ↓
SIM ↑
Mel ℓ1
15M
3.79
68.42
5.41
2.76
3.55
2.25
0.883
Mel ℓ1
30M
3.70
69.28
5.46
2.73
3.61
2.22
0.883
Full-mel FM
30M
3.72
68.92
5.59
2.61
3.58
2.26
0.875
SPEAR-Gen
30M
3.71
69.48
5.08
2.94
3.69
2.22
0.892
Table 5: Ablation of the generative objectives. #Params denotes the total number of parameters in decoders.
Integrating speech understanding and generation is a pivotal step toward building unified speech models. However, the different representations required for these two tasks currently pose significant compatibility challenges. Typically, semantics-oriented features are learned from self-supervised learning (SSL), and acoustic-oriented features from reconstruction. Such fragmented representations hinder the realization of truly unified speech systems. We present WavCube, a compact continuous latent derived from an SSL speech encoder that simultaneously supports speech understanding, reconstruction, and generation. WavCube employs a two-stage training scheme. Stage 1 trains a semantic bottleneck to filter off-manifold redundancy that makes raw SSL features intractable for diffusion. Stage 2 injects fine-grained acoustic details via end-to-end reconstruction, while a semantic anchoring loss ensures the representation remains grounded within its original semantic manifold. Comprehensive experiments show that WavCube closely approaches WavLM performance on SUPERB despite an 8x dimensional compression, attains reconstruction quality on par with existing acoustic representations, delivers state-of-the-art zero-shot TTS performance with markedly faster training convergence, and excels in speech enhancement, separation, and voice conversion tasks on the SUPERB-SG benchmark. Systematic ablations reveal that WavCube's two-stage recipe resolves two intrinsic flaws of SSL features for generative modeling, paving the way for future unified speech systems. Codes and checkpoints are available at https://github.com/yanghaha0908/WavCube.
Guanrou Yang, Tian Tan, Qian Chen +12
Shanghai Jiao Tong University · Independent Researcher · Shanghai Innovation Institute +4
Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems that fail to capture speaker-specific acoustic variation. Prior work demonstrates that grapheme-based models outperform phoneme-based systems at scale, but not in low-resource settings. In this paper, we propose SPARCLE, a speaker-aware grapheme representation model that enriches characters with their precise acoustic realizations. SPARCLE is trained with a contrastive objective to align graphemes with corresponding Wav2Vec2 acoustic representations while conditioned on speaker identity. The resulting model serves as a replacement to G2P systems for downstream text-to-speech (TTS) tasks. We demonstrate that SPARCLE improves generation quality, reducing word error rates by half in extreme low-resource settings compared to standard grapheme-based models.
Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2
University of Illinois Urbana-Champaign, USA · National Center for Supercomputing Applications, USA
We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives. OLIVE combines view-augmented masked latent prediction with waveform reconstruction under a unified objective. Reconstruction constrains early encoder features to retain signal-level information, while masked latent prediction shapes later contextual representations toward invariance for robust downstream performance. We show that these objectives enable representations that support a broad range of tasks. In particular, OLIVE improves results on generation and speaker tasks, maintains competitive performance on recognition and semantic tasks, and improves waveform reconstruction.
Karl El Hajal, Mathew Magimai. -Doss
Idiap Research Institute, Switzerland · EPFL, Switzerland