cs.CLSep 13, 2026

Quantifying the Generation Modality Gap in Speech-Text Language Models

Authors: Ju-Chieh Chou, Jiawei Zhou, Karen Livescu

Organizations: TTI-Chicago · Stony Brook University

Abstract

Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and scoring it with a reference language model; local phonetic structure, measured by phone n-gram distributional statistics; speaker consistency and acoustic quality; and emotion-based distributional metrics. Across datasets, we find that joint speech-text modeling substantially improves semantic coherence. However, the improvement is not uniform across metrics: phone-level metrics change only modestly, speaker similarity and predicted quality are lower for speech-text continuations, while emotion-based distributional metrics improve. Compared with larger-scale speech-only models, our speech-text model closes much of the scaling gap in transcript-based semantic coherence, suggesting that text provides an efficient semantic training signal for spoken language modeling.

Figures & tables

Explore similar work

CardsList
  1. Why We Need Speech to Evaluate Speech Translation

    May 27, 2026Maike Züfle, Danni Liu, Vilém Zouhar +1Speech TranslationSpeaker

  2. Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

    May 7, 2026Wenqian Cui, Xiao-Hui Li, Daxin Tan +2Speech Language ModelsModality Gap

  3. MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

    Aug 12, 2026Xingwei Sun, Heinrich Dinkel, Gang Li +7Text-To-AudioSeed-Tts-Eval Benchmark