LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech remains unclear. We introduce two emotion-sensitive metrics for matched neutral and emotional syntheses---a codec trajectory score and a late residual direction score---and use them to score activation-patching interventions. Under controlled matched-reference conditions, this analysis identifies a sparse source-to-readout component-level circuit: 23--27 attention heads and MLPs per emotion, roughly 5% of the components considered, recover or suppress 74--88% of the late emotion-readout shift on held-out cases. The circuit combines a shared component backbone with emotion-specific components; cross-emotion activation swaps reduce the target readout in 47 of 48 cases. In decoded speech, the same intervention produces consistent changes in pitch, energy, and spectral brightness over 24 matched pairs per emotion. A readout-matched residual-direction baseline produces only 17--27% of the intervention's pitch effect, showing that internal readout movement alone does not explain the decoded acoustic changes. These results trace a compact causal route from reference-derived prefix information to emotion-relevant properties of generated speech.
Figures & tables
Figure 1: Causal tracing localizes a sparse source-to-readout component-level emotion circuit within a reference-conditioned LLM-based TTS model. Matched neutral/emotional runs localize a conditioning-prefix source, trace its effect through sparse attention-head and MLP sets to late emotion readouts, and validate the pathway with causal and decoded-speech tests.
Emotion
Codec pair
Codec split
Resid. cent.
Anger
0.941
0.998
0.926
Happiness
0.764
0.988
0.903
Sadness
0.903
0.996
0.874
Table 1: Stability of the two emotion metrics over 50 matched pairs. For the codec trajectory score, Codec pair reports the mean pairwise cosine among paired codec-token embedding shifts, and Codec split reports split-half agreement of the global codec direction. Resid. cent. reports the mean cosine to the final-layer residual centroid used by the late residual direction score.
Figure 2: Conditioning-prefix source localization and direct source tests. (a) Layer-by-position scans show that emotional-prefix patching into neutral runs has the strongest final-layer residual recovery in the conditioning prefix, with a narrow peak near positions [12,16) . (b) Direct interventions on the layer-0 [0,32) source window validate this source across anger, happiness, and sadness: emotional-prefix patching recovers the emotion effect, while neutral-prefix patching removes it. Ratios near one indicate recovery or removal of nearly the full clean neutral–emotion shift under both codec trajectory and late residual direction scores.
Emotion
Nodes
Heads
MLPs
Late MLPs
Shared
Anger
24
15
9
6
13
Happiness
23
13
10
6
13
Sadness
27
21
6
5
13
Table 2: Gradient-attribution candidate sets. Nodes are selected attention-head and MLP components in the frozen top-30% candidate set for each emotion. The Late MLPs column counts selected MLPs in layers 18–23. Shared counts components that appear in all three emotion-specific candidate sets.
Emotion
Components
Pilot recover
Pilot remove
Valid. recover
Valid. remove
Anger
24 (15/9)
0.812 / 0.815
0.787 / 0.786
0.814 / 0.815
0.784 / 0.787
Happiness
23 (13/10)
0.883 / 0.886
0.864 / 0.865
0.876 / 0.868
0.855 / 0.855
Sadness
27 (21/6)
0.751 / 0.754
0.754 / 0.755
0.736 / 0.739
0.747 / 0.744
Table 3: Pilot and validation results for frozen top-30% component sets. Each component set is selected on eight pilot pairs and evaluated without retuning on eight validation synthesis pairs. Components reports total selected components, with heads/MLPs in parentheses. Recover is the fraction of the full-source readout shift restored by the selected components; remove is the fraction lost when the same components are replaced with neutral activations in the full-source trajectory. Each recover/remove cell reports mean/median.
Emotion
Recover
Remove
Selected
Random
Selected
Random
Anger
0.814
0.278
0.784
0.242
Happiness
0.876
0.421
0.855
0.463
Sadness
0.736
0.340
0.747
0.395
Table 4: Matched random controls for component-set validation. Selected cells report validation mean ratios for the frozen component sets. Random cells report the mean over 20 matched random sets with the same size and head–MLP mix. The selected set is larger than every matched random set in all three emotions and both intervention directions.
Emotion
Full
Shared only
Specific only
Anger
0.814
0.686 (84%)
0.470 (58%)
Happiness
0.876
0.753 (86%)
0.511 (58%)
Sadness
0.736
0.584 (79%)
0.433 (59%)
Table 5: Sufficiency recovery of full, shared-only, and emotion-specific component sets. Parentheses give the fraction of the full-set recovery.
Emotion
Circuit Δ F0 (st)
Second correlate
Circuit change
Residual Δ F0
Ratio
Anger
+2.77[2.44,3.12]
RMS energy
+3.35 dB [2.99,3.73]
+0.74 st
0.27
Happiness
+3.57[3.18,3.98]
Spectral centroid
+133.06 Hz [113.37,152.10]
+0.83 st
0.23
Sadness
−1.01[−1.38,−0.55]
RMS energy
−3.04 dB [−3.42,−2.67]
−0.17 st
0.17
Table 6: Decoded acoustic effects of the sparse component intervention and a readout-matched residual-direction baseline. Changes are measured from the matched neutral baseline. The second correlate follows the salient target-reference contrast for each emotion.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Removed
Pair cos.
Split cos.
0
0.764
0.988
1
0.769
0.986
3
0.780
0.988
5
0.789
0.987
10
0.810
0.987
Appendix
Table 7: Outlier removal for the Happiness codec direction. Removing the least aligned happiness pairs modestly increases pairwise agreement but leaves split-half agreement high.
Figure 3: Emotion clustering at the first autoregressive prediction step. Each point is a paired sample-level emotion direction, computed as the residual difference between an emotional trajectory and its neutral counterpart at the same layer and prediction step. t-SNE is fit independently within each layer. This figure measures layerwise separability at a single prediction step; the whole-generation residual readout used for component selection and validation is computed separately.
Emotion
Mean [95% CI]
Wins
p
Anger
+0.323[−0.209,0.833]
17/24
0.243
Happiness
+0.730[0.438,1.047]
20/24
2×10−5
Sadness
+0.396[0.105,0.731]
19/24
0.0156
Appendix
Table 8: Paired statistics for calibrated Emo-SIM change.
Emotion
Δ F0 in semitones [95% CI]
Anger
+2.15[1.29,3.00]
Happiness
+1.69[1.20,2.22]
Sadness
−0.95[−1.46,−0.45]
Appendix
Table 9: F0 effects under top- p decoding. Intervals are paired bootstrap 95% confidence intervals over eight pairs per emotion.
Emotion
Emotion pref.
p
Naturalness
Anger
8/8
0.0039
0/1/7
Happiness
8/8
0.0039
1/1/6
Sadness
8/8
0.0039
0/0/8
Overall
24/24
5.96×10−8
1/2/21
Appendix
Table 10: Single-listener blinded A/B results. Naturalness cells report intervention/baseline/tie counts. The reported p -values are item-level one-sided exact sign tests.
Emotion
Codec suff.
Resid. suff.
Codec nec.
Resid. nec.
Anger
0.62
0.60
0.47
0.54
Happiness
0.55
0.55
0.73
0.65
Sadness
0.59
0.53
0.59
0.64
Appendix
Table 11: Direct tests on the narrow prefix hotspot [12,16) . Ratios are averaged over 50 matched pairs per emotion. The hotspot recovers or removes a substantial fraction of the clean neutral–emotion shift, but does not fully account for the effect; the main tracing experiments therefore use the broader layer-0 [0,32) conditioning-prefix block as the source window.
Anger
Happiness
Sadness
L21MLP
L22MLP
L0H2
L22MLP
L21MLP
L21MLP
L20MLP
L0H2
L22MLP
L18MLP
L20MLP
L20MLP
L16H8
L19MLP
L0MLP
L17MLP
L6H2
L0H16
Appendix
Table 12: Frozen canonical component identities. Each column follows its emotion-specific attribution ranking. Bold entries are the 13 components shared by all three sets.
Figure 4: Source-conditioned component specificity controls. Each selected component is compared with five cleaner matched alternatives: same-layer non-selected heads for head components, and nearby same-type non-selected MLPs for MLP components. Bars report pair-level selected-minus-control deltas after averaging over selected components within each validation pair; error bars are bootstrap confidence intervals over the eight validation pairs. Positive values indicate that selected components carry more source-conditioned readout effect than matched alternatives.
Target
Mean change
95% CI
Lower
Anger
−33.14
[−47.44,−18.84]
15/16
Happiness
−40.72
[−52.73,−28.70]
16/16
Sadness
−51.46
[−53.47,−49.46]
16/16
Overall
−41.77
[−48.27,−35.28]
47/48
Appendix
Table 13: Cross-emotion shared-component swaps. Negative values mean that another emotion’s donor activations weaken the target readout.
Integrating large language models (LLMs) into text-to-speech (TTS) systems has improved speech expressiveness, yet interpretable emotional control remains challenging. Existing approaches primarily rely on external conditioning or global activation steering, offering limited insight into the internal representations underlying emotional control. In this work, we analyze emotion-related variation in the semantic hidden states of LLM-based TTS models using sparse autoencoders (SAEs) to identify sparse latent features. Our analysis shows that emotional variation is distributed across multiple sparse latent features, while intervening on a small subset enables interpretable emotion control. Building on this observation, we introduce a feature-level intervention framework for bidirectional emotion induction and suppression without modifying backbone parameters. We further show that distinct latent features are associated with specific acoustic attributes (e.g., pitch), suggesting that emotional expression arises from coordinated latent contributions rather than a single global shift. Empirically, steering these sparse latent features achieves comparable or superior emotion induction and suppression performance relative to global steering and existing TTS baselines.
Hongfei Du, Jiacheng Shi, Sidi Lu +2
Department of Computer Science, William & Mary, USA.
Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic -- acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as a learning-to-rank problem and propose Emo-LiPO, a listwise preference optimization framework that aligns prompt-conditioned speech generation with relative emotion intensity expressed in text. Emo-LiPO explicitly models global intensity ordering within each emotion under fixed transcripts, enabling more faithful and continuous emotional expression. We further construct ESD-plus, a multi-speaker dataset with explicit emotion intensity variations, to support fine-grained emotion modeling and evaluation. Experiments on ESD-plus demonstrate that Emo-LiPO significantly improves emotion accuracy and intensity controllability over both supervised- and DPO-based LLM TTS baselines, with particularly pronounced gains at high intensity levels.
Yihang Lin, Li Zhou, Congwei Cao +4
The Chinese University of Hong Kong, Shenzhen · Shenzhen Loop Area Institute · Agency for Science, Technology and Research +2
SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.