cs.SDSep 24, 2026

Speech Block Influence: Component-Specific Layer Scoring for Pruning Speech LLMs

Authors: Siyu Yao, Du Q. Huynh, Lian Xu, Mark Reynolds

Organizations: The University of Western Australia

Abstract

Speech LLMs are costly to deploy in resource-constrained settings. Layer pruning can cut this cost, but existing scoring metrics transfer poorly to speech LLMs: they assume a decoder-only architecture with homogeneous token sequences, whereas speech LLMs add encoder and adapter components and process multimodal sequences. We propose Speech Block Influence (SBI), the first layer-importance scoring framework designed for speech LLM pruning that consists of two component-specific scores: SBI-Enc measures the effect of encoder-layer removal at the adapter's output to better reflect downstream impact; SBI-Dec measures layer-wise input-output similarity over text-token positions only to avoid audio-token dominance. Across three speech LLMs, SBI improves pruning robustness, with stronger encoder performance at higher pruning rates and more reliable decoder layer selection by scoring text tokens rather than the audio-dominated full sequence. We further find that text-only calibration yields decoder rankings highly correlated with those from speech-text calibration, suggesting a cheaper alternative to measure decoder layer importance.

Figures & tables

Explore similar work

CardsList
  1. Measuring the Redundancy of Decoder Layers in SpeechLLMs

    Mar 5, 2026Adel Moumen, Guangzhi Sun, Philip C WoodlandLarge Language Model DecodingDecoder Layers

  2. Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

    Jul 7, 2026Ke-Han Lu, Keqi Deng, Ruchao Fan +2Speech Language ModelsAutoregressive Decoding

  3. Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs

    Sep 29, 2026Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi ShekharCharacter Error RateDisparities