From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS
Organizations: Independent Researcher
Abstract
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.
Figures & tables
| Model | Speaker similarity | Instruction following | ||||
|---|---|---|---|---|---|---|
| ZH | EN | Yes% | Partial% | No% | Avg. score | |
| Proposed system and ablations | ||||||
| ControlEdit-TTS (full) | 0.776 | 0.791 | 95.5 | 4.2 | 0.3 | 4.819 |
| ControlEdit-TTS (edit disabled) | 0.777 | 0.791 | 87.0 | 12.3 | 0.7 | 4.639 |
| ControlEdit-TTS (one-shot) | 0.752 | 0.776 | 85.2 | 13.8 | 1.0 | 4.592 |
| Agentic baselines with our framework | ||||||
| Opponent | Chinese ( ) | English ( ) | ||
|---|---|---|---|---|
| W/L/T | WR | W/L/T | WR | |
| Direct dialogue systems | ||||
| SoulX-Podcast | 81/12/0 | 0.871 | 71/19/4 | 0.755 |
| Fish Audio S2 | 73/19/1 | 0.785 | 67/18/9 | 0.713 |
| VibeVoice-7B | 56/36/1 | 0.602 | 59/26/9 | 0.628 |
| VibeVoice-1.5B | 76/17/0 | 0.817 | 63/19/12 | 0.670 |