cs.CLOct 8, 2026

Local Prototype Reconstruction for Text-Compatible Speech-to-LLM Bridge Pretraining

Authors: Xinnian Zhao, Chia-Hua Wu, Pu Wang, Hugo Van Hamme

Organizations: Department of Electrical Engineering (ESAT), KU Leuven, Leuven, Belgium · Institute of Information Science, Academia Sinica, Taiwan

Abstract

Speech-to-LLM systems often connect a frozen speech encoder to a frozen large language model (LLM) through a small trainable bridge. The bridge is usually treated as plumbing, but it in fact defines the geometry of the speech-to-LLM interface, and the pretraining objective decides whether that interface provides a reusable initialization for downstream tasks. We study a transferable bridge through two complementary properties: global alignment with the text side, and local lexical manifold compatibility, where bridge embeddings remain close to the frozen LLM's input-embedding neighbourhoods. We make this property measurable with a fixed, head-free, timestamp-free diagnostic that applies to any objective, and show that next-word prediction (NWP) and sentence-level contrastive pretraining do not fully capture token-level lexical compatibility. We then introduce Local Prototype Reconstruction (LPR), a lightweight training-only regularizer that requires each aligned bridge token to be reconstructable from a small neighbourhood of frozen LLM token embeddings, with a hard single-prototype anchor as its limiting case. On multilingual ASR and speech translation, LPR improves transfer, with the largest gains on translation and low-resource adaptation. Crucially, our independent diagnostic correlates with downstream gains across objectives, suggesting that lexical manifold compatibility is predictive of reusability for speech-to-LLM bridges.

Figures & tables

Explore similar work

CardsList
  1. Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

    Jun 8, 2026Ming-Hao Hsu, Yuxuan Hu, Shujie Liu +3Audio Representation LearningAudio Understanding

  2. DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs

    Oct 6, 2026Hemant Yadav, Sunayana Sitaram, Roger Zimmermann +1Automatic Speech RecognitionSpeech Language Models