Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.
Figures & tables
Figure 1: Overview of our method.
Figure 2: Overview of the proposed method. (a) The agentic refinement framework converts an input script into a structured dialogue plan and iteratively applies utterance- and scene-level critique to trigger attribute editing, selective resynthesis, or timing adjustment. (b) ControlEdit-TTS unifies instruction-following synthesis and attribute-level editing within a shared autoregressive model, enabling targeted correction of emotion, speaking rate, and energy/loudness while preserving linguistic content and speaker identity.
Model
Speaker similarity ↑
Instruction following
ZH
EN
Yes% ↑
Partial%
No% ↓
Avg. score ↑
Proposed system and ablations
ControlEdit-TTS (full)
0.776
0.791
95.5
4.2
0.3
4.819
ControlEdit-TTS (edit disabled)
0.777
0.791
87.0
12.3
0.7
4.639
ControlEdit-TTS (one-shot)
0.752
0.776
85.2
13.8
1.0
4.592
Agentic baselines with our framework
Table 1: Utterance-level instruction following and speaker preservation. The automatic judge assigns Yes, Partial, or No and a score from 1 to 5. One-shot denotes the shared initial synthesis before critique-guided refinement.
Opponent
Chinese ( N=93 )
English ( N=94 )
W/L/T
WR
W/L/T
WR
Direct dialogue systems
SoulX-Podcast
81/12/0
0.871
71/19/4
0.755
Fish Audio S2
73/19/1
0.785
67/18/9
0.713
VibeVoice-7B
56/36/1
0.602
59/26/9
0.628
VibeVoice-1.5B
76/17/0
0.817
63/19/12
0.670
Table 2: Gemini pairwise preference for ControlEdit-TTS (full). W/L/T denote our wins/losses/ties, and WR=W/(W+L+T) . The utterance-only variant omits scene-level critique and refinement.
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to transcript spans or boundaries. This semantic timeline avoids explicit timestamp alignment and provides an externally inspectable contract for compositional edits. We instantiate the interface in dots.tts.edit, an editor adapted from the continuous autoregressive dots.tts foundation model. Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing through text, emotion, prosody, and pause editing. Task-specific data pipelines construct operation- and scope-controlled pairs while retaining source-derived context outside each target region. We further introduce doteBench, a bilingual evaluation suite that measures precise instruction following, local preservation, and audio quality across the four controls and their composition. Experiments show leading overall instruction following and local preservation across its five editing categories, while audio quality remains comparable to existing open-source systems. Across three Seed-TTS-Eval shards, the model shows negligible differences from the base model in zero-shot TTS recognition error rate and speaker similarity.
Hankun Wang, Bohan Li, Shi Lian +7
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University · Xiaohongshu Inc.
Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.
Sihang Nie, Xueru Li, Xiaofen Xing +4
South China University of Technology · Huya Inc. · The Hongkong Polytechnic University
Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.