Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.
Figures & tables
Figure 1: TopK SAE produce zt on top of a frozen speech encoder, which is routed into linguistic zL and paralinguistic zP representations.
LibriSpeech
MSP-Podcast
Encoder
SAE Model
Route
Recon. ↓
PER ↓
SID Acc. ↑
Recon. ↓
PER ↓
SID Acc. ↑
Emo. UAR ↑
F0 corr. ↑
Energy corr. ↑
SPEAR X-Large
Recon-only
zt
0.066
6.00
99.6
0.075
32.37
61.74
65.25
0.578
0.944
Fixed
zt
0.179
10.50
100.0
0.146
30.52
65.98
66.25
0.482
0.923
zL
10.02
6.2
30.58
20.67
51.71
0.118
0.781
zP
99.60
99.8
81.79
76.47
67.40
0.564
0.918
Quota-freeze
zt
0.163
8.73
100.0
0.171
34.74
68.16
67.46
0.485
0.924
Table 1: Independent probes across encoders and corpora. Reconstruction MSE (Recon.) is reported once per SAE model.
Outcome
Fixed
Freeze
P -swap phone retention
99.57
99.16
P -swap donor-speaker match
98.00
97.20
L -swap phone replacement
94.26
94.12
L -swap recipient-speaker match
96.00
96.80
L -swap donor-phone retention
99.10
98.84
Table 2: Controlled route interventions on the fixed 250-pair SPEAR evaluation set (%).
Routing
Condition
zL PER
zP PER
zL SID
zP SID
(%)
(%)
(%)
(%)
Fixed
Full model
10.02
99.60
6.2
98.6
No adversaries
4.49
66.64
99.8
100.0
No speaker adv.
4.63
85.25
100.0
99.4
No phone adv.
11.01
47.82
1.6
100.0
Freeze
Full model
8.94
80.26
0.6
99.6
Table 3: Ablations of the adversarial objectives on SPEAR. Lower zL SID and higher zP PER indicate stronger cross-factor suppression.
Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of paralinguistic information in four open source models, Whisper-large-v2, Qwen2-Audio-7B Instruct, Qwen2.5-Omni-7B, and Chroma-4B, using the Expresso dataset with controlled speaking styles. We combine centered kernel alignment, linear probing with leave one speaker out evaluation, open ended tone prediction, and a content prosody leakage metric to trace how style information moves from the audio encoder to the final output. All models strongly encode speaking style in the late encoder, that is, the top third of the audio encoder's layers, but this information is consistently degraded before reaching the output. The projector reshapes representation geometry without removing information, while decoders differ in how much style they preserve depending on architecture and training objective. At the output level, models fall into two behaviors. Some are content driven, where predictions depend mainly on text. Others are acoustic driven, where predictions vary with speaking style. The leakage metric quantifies this difference, and qualitative results confirm it. Overall, we identify a gap between what models encode and what they use, highlighting a key limitation in current audio language models.
Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh +1
Language Technologies Institute Carnegie Mellon University
Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream. We train BatchTopK sparse autoencoders on the LM backbone of CosyVoice3 and introduce a modality-aware auto-interp pipeline that labels each feature from where it fires-text-prefix context, 1-second speech clips, or both. The recovered features are interpretable, spanning phonemes, laughter, accent prompts and speaker gender. Steering through the SAE latent space shows these features are causal rather than merely descriptive: targeted interventions raise laughter probability from 0.02 to 0.79, flip perceived speaker gender, and control speech rate while preserving spoken content. SAE features thus serve both as interpretability objects and as control directions for TTS synthesis.
Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1
While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understanding text-based transformer models, leaving ASR systems largely unexplored. In order to address this gap, we examine the internal representations of Whisper's encoder using a sparse autoencoder. We find diverse monosemantic features across linguistic and non-linguistic boundaries, spanning a hierarchy from phonetic to semantic representations, and conduct a causal feature-steering campaign across this hierarchy, including cross-lingual steering. We further find that steering is more reliable for higher-level features than lower-level ones, an asymmetry that may reflect redundant encoding of lower-level information. Altogether, this work demonstrates that Whisper's encoder represents a surprisingly rich hierarchy of linguistic information that extends well beyond what is strictly necessary for transcription.