cs.SDOct 4, 2026

Tracing a Sparse Emotion-Control Circuit in LLM-Based Text-to-Speech

Authors: Hongfei Du, Jiacheng Shi, Yanfu Zhang, Ye Gao

Organizations: Department of Computer Science, William & Mary, Williamsburg, USA

Abstract

LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech remains unclear. We introduce two emotion-sensitive metrics for matched neutral and emotional syntheses---a codec trajectory score and a late residual direction score---and use them to score activation-patching interventions. Under controlled matched-reference conditions, this analysis identifies a sparse source-to-readout component-level circuit: 23--27 attention heads and MLPs per emotion, roughly 5% of the components considered, recover or suppress 74--88% of the late emotion-readout shift on held-out cases. The circuit combines a shared component backbone with emotion-specific components; cross-emotion activation swaps reduce the target readout in 47 of 48 cases. In decoded speech, the same intervention produces consistent changes in pitch, energy, and spectral brightness over 24 matched pairs per emotion. A readout-matched residual-direction baseline produces only 17--27% of the intervention's pitch effect, showing that internal readout movement alone does not explain the decoded acoustic changes. These results trace a compact causal route from reference-derived prefix information to emotion-relevant properties of generated speech.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

    May 31, 2026Hongfei Du, Jiacheng Shi, Sidi Lu +2Emotional RegulationExpressivity

  2. Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

    Jun 11, 2026Yihang Lin, Li Zhou, Congwei Cao +4Emotion Recognition

  3. Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

    Sep 17, 2026Hasindri Watawana, Sergio Burdisso, Esaú Villatoro-Tello +4Emotion RecognitionSpeech Language Models