cs.CLOct 7, 2026

Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders

Authors: Beimnet Bekele Guta, Xiaoyu Yang, Guangzhi Sun, Philip C. Woodland

Organizations: Department of Engineering, University of Cambridge, Cambridge, UK

Abstract

Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.

Figures & tables

Explore similar work

CardsList
  1. Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

    Sep 1, 2026Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh +1Audio-Language Model EvaluationAudio-Language Models

  2. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    Jun 8, 2026Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1TTS SynthesisLanguage Model Steering

  3. On the Interpretability of Whisper Encodings Using Sparse Autoencoders

    May 12, 2026Dan Pluth, Zachary Nicholas Houghton, Yu Zhou +1Transformer InterpretabilityMechanistic Interpretability