cs.SDSep 27, 2026

Controlling Speaking Rate in Autoregressive TTS via Activation Steering

Authors: Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet, Juan Pablo Zuluaga Gomez

Organizations: AGIGO · Sapienza University of Rome · ETH Zurich · University of Zurich

Abstract

Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.

Figures & tables

Explore similar work

CardsList
  1. TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

    Jun 8, 2026Yejin Lee, Junwon Moon, Hyoeun Kim +3Autoregressive Text-To-SpeechAutoregressive Decoding

  2. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

    Jul 5, 2026Junwon Moon, Seungbeom Kim, Yejin Lee +4Autoregressive Text-To-SpeechDiffusion Language Models