eess.ASSep 28, 2026

SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations

Authors: Xiaoyu Yang, Arthur Hinsvark, Antonios Alexos, Osama Hanna, Philip C. Woodland, Yiting Lu

Organizations: Department of Engineering, University of Cambridge, Cambridge, UK · Meta Superintelligence Labs, Menlo Park, USA

Abstract

Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.

Figures & tables

Explore similar work

CardsList
  1. WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

    May 7, 2026Guanrou Yang, Tian Tan, Qian Chen +12Flow-Matching Text-To-Speech

  2. SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

    May 1, 2026Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2Discrete Speech RepresentationsGrapheme-To-Phoneme

  3. OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

    Jun 29, 2026Karl El Hajal, Mathew Magimai. -DossSelf-Supervised Speech ModelsSelf-Supervised Learning