cs.CLOct 6, 2026

Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model

Authors: Pulak Kuli

Organizations: Independent Researcher

Abstract

Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Steerspeech: Activation Steering For Emotion Control In Generated Speech

    Oct 7, 2026Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1

  2. SteerDuplex: Steerable Duplex Speech Dialogue Models

    Sep 14, 2026Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai +13Full-Duplex Speech ModelsTurn-Taking

  3. A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

    Jul 1, 2026Siyi Wang, James Bailey, Ting DangSpeech Language ModelsEmotion Recognition