Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.
Figures & tables
Figure 1: Motivation and overview of InstCharVoice. The framework bridges natural-language instructions and explicit character-level acoustic control through instruction grounding and acoustic planning.
Figure 2: Model architecture of InstCharVoice, including (a) training sequence construction, (b) autoregressive modeling of character-level speech chunks, and (c) grounding-aware training objective.
Statistic
Utterances
Keywords
Dur
Bnd
Eng
Pit
Ton
Count
2143k
5978k
609k
3882k
2830k
202k
901k
Table 1: Statistics of the character-level grounded instruction annotations. Dur, Bnd, Eng, Pit, and Ton denote duration, boundary, energy, pitch, and tone, respectively.
Model
DNSMOS ↑
CER ↓
kD-MAE ↓
kE-MAE ↓
kP-MAE ↓
kB-RER ↓
kT-RER ↓
GroundTruth
3.616
—
0.0182
0.0133
0.0094
5.26%
8.95%
FireRedTTS3
3.599
2.89%
0.1162
0.1409
0.2326
31.29%
47.02%
Qwen3-TTS
3.642
2.99%
0.1057
0.1930
0.2291
28.36%
44.56%
CosyVoice3
3.583
3.12%
0.1126
0.1411
0.2383
24.27%
47.84%
InstCharVoice
3.686
3.58%
0.0769
0.1271
0.1857
14.04%
37.05%
- w/o Instruction
3.560
2.91%
0.0965
0.1431
0.2401
23.10%
44.08%
Table 2: Objective comparison of different systems.
Figure 3: Blind evaluation of instruction annotation consistency.
While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the structural mismatch between discrete textual intents and continuous acoustic realizations. Inspired by human cognitive decoupling, we introduce AgentSteerTTS, a multi-agent closed-loop framework designed for intent-faithful expressive control of composite instructions. First, in our framework, an adversarial disentanglement agent mitigates speaker-emotion leakage by learning separable identity and emotion-prosody subspaces with leakage-suppressing regularization. Next, a Dual-Stream Anchoring Controller grounds abstract intents using a large-scale acoustic prototype library: a Retrieval Agent selects expressive anchors, while a Synthesis Agent fuses them into continuous control vectors via gated attention. Finally, a Fast-Slow Feedback Agent refines output intensity through latent gradient correction and resolves semantic-acoustic mismatches using high-level perceptual critique. Experiments on a composite-instruction benchmark and public test sets show that AgentSteerTTS yields consistent and significant improvements to the baselines, demonstrating the effectiveness of the proposed method.
Bin Kang, Shaoguo Wen, Yang Fan +6
University of Chinese Academy of Sciences · Shenzhen Loop Area Institute · Tencent Turinglab +2
While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In scenarios demanding precise stylistic interventions and strict temporal alignment, such as audiobook narration and video dubbing, the inability to explicitly manipulate word-level acoustic attributes remains a critical bottleneck. This limitation is primarily amplified by the severe scarcity of fine-grained annotated datasets and the architectural challenge of integrating multi-dimensional control signals into discrete autoregressive generation. To address this, we propose a unified framework for highly precise word-level control. First, we construct WordVoice-5A, a massive 4.7k-hour bilingual dataset featuring five-dimensional word-level annotations (duration, boundary, energy, pitch and tone) developed through a rigorous linguistically-guided pipeline. Second, we introduce WordVoice to transform the implicit generation process into an explicit, highly controllable paradigm. Specifically, we introduce a bound-token mechanism within the LLM to formulate an explicit ``acoustic planning'' process, enabling adaptive multi-task prosodic planning and flexible manual intervention. Furthermore, we augment the token-to-waveform stage with a fine-grained acoustic modulation module, bridging the resolution gap to strictly align word-level attributes between highly compressed discrete tokens and continuous waveforms. Extensive experiments demonstrate that WordVoice achieves superior, decoupled control over multiple acoustic dimensions while maintaining competitive zero-shot synthesis stability. The code and audio samples are publicly available at https://xxh333.github.io/wordvoice-demo/.
Sihang Nie, Jinxin Ji, Xiaofen Xing +4
South China University of Technology · Huya Inc. · Tongji University +2
Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity, and insufficient multilingual support. We present \textbf{MINT-Bench}, a comprehensive multilingual benchmark for instruction-following TTS. MINT-Bench is built upon a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol that jointly assesses content consistency, instruction following, and perceptual quality. Experiments across ten languages show that current systems remain far from solved: frontier commercial systems lead overall, while leading open-source models become highly competitive and can even outperform commercial counterparts in localized settings such as Chinese. The benchmark further reveals that harder compositional and paralinguistic controls remain major bottlenecks for current systems. We release MINT-Bench together with the data construction and evaluation toolkit to support future research on controllable, multilingual, and diagnostically grounded TTS evaluation. The leaderboard and demo are available at https://longwaytog0.github.io/MINT-Bench/
Huakang Chen, Jingbin Hu, Liumeng Xue +12
School of Intelligence Science and Technology, Nanjing University, Suzhou, China · Yutu Zhineng, Beijing, China · Lingguang Zhaxian Technology, Shanghai, China