cs.SDOct 8, 2026

Edit Who Speaks, Control How They Speak: Global Timbre Editing and Local Instruction Control for TTS

Authors: Junchuan Zhao, Chenglin Xu, Wei Zeng, Haoyang Li, Yiwen Guo, Ye Wang

Organizations: National University of Singapore · LIGHTSPEED · Nanyang Technological University · Independent Researcher

Abstract

Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce \textbf{EDICT}, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality. Audio demos are available.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

    Aug 2, 2026Hankun Wang, Bohan Li, Shi Lian +7Speech Foundation ModelsSpeech Editing

  2. CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

    May 25, 2026Junyang Chen, Yuhang Jia, Hui Wang +3Speech EditingZero-Shot TTS

  3. EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows

    Sep 24, 2026Hongyao Deng, Wenhao Guan, Xuetao Lin +4TTS SynthesisSpeech Editing