Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.
Figures & tables
Figure 1: Motivation and overview of InstCharVoice. The framework bridges natural-language instructions and explicit character-level acoustic control through instruction grounding and acoustic planning.
Figure 2: Model architecture of InstCharVoice, including (a) training sequence construction, (b) autoregressive modeling of character-level speech chunks, and (c) grounding-aware training objective.
Statistic
Utterances
Keywords
Dur
Bnd
Eng
Pit
Ton
Count
2143k
5978k
609k
3882k
2830k
202k
901k
Table 1: Statistics of the character-level grounded instruction annotations. Dur, Bnd, Eng, Pit, and Ton denote duration, boundary, energy, pitch, and tone, respectively.
Model
DNSMOS ↑
CER ↓
kD-MAE ↓
kE-MAE ↓
kP-MAE ↓
kB-RER ↓
kT-RER ↓
GroundTruth
3.616
—
0.0182
0.0133
0.0094
5.26%
8.95%
FireRedTTS3
3.599
2.89%
0.1162
0.1409
0.2326
31.29%
47.02%
Qwen3-TTS
3.642
2.99%
0.1057
0.1930
0.2291
28.36%
44.56%
CosyVoice3
3.583
3.12%
0.1126
0.1411
0.2383
24.27%
47.84%
InstCharVoice
3.686
3.58%
0.0769
0.1271
0.1857
14.04%
37.05%
- w/o Instruction
3.560
2.91%
0.0965
0.1431
0.2401
23.10%
44.08%
Table 2: Objective comparison of different systems.
Figure 3: Blind evaluation of instruction annotation consistency.
School of Intelligence Science and Technology, Nanjing University, Suzhou, China · Yutu Zhineng, Beijing, China · Lingguang Zhaxian Technology, Shanghai, China