eess.ASSep 28, 2026

InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech

Authors: Sihang Nie, Xueru Li, Xiaofen Xing, Deyi Tuo, Cheng-Bin Jin, Jingyuan Xing, Jinxin Ji

Organizations: South China University of Technology · Huya Inc. · The Hongkong Polytechnic University

Abstract

Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.

Figures & tables

Explore similar work

CardsList
  1. AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

    May 14, 2026Bin Kang, Shaoguo Wen, Yang Fan +6Text-To-Speech SynthesisAcoustic Latent Space

  2. WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

    Jul 7, 2026Sihang Nie, Jinxin Ji, Xiaofen Xing +4Flow-Matching Text-To-SpeechProsody

  3. MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

    Apr 20, 2026Huakang Chen, Jingbin Hu, Liumeng Xue +12Seed-Tts-Eval BenchmarkMultilingual Benchmark