Controllable Speech Generation

Momentum

21 papers in the last four weeks, up 600% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 84

All topics
CardsList
  1. dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

    Aug 2, 2026Hankun Wang, Bohan Li, Shi Lian +7Speech Foundation ModelsSpeech Editing

  2. Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

    Aug 1, 2026Xianhao Zhou, Jianghao WuSpeech Generation EvaluationTTS Evaluation

  3. Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

    Jul 20, 2026Shengfan Shen, Di Wu, Xingchen Song +5TTS SynthesisLanguage Model-Based Control

  4. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

    Jul 14, 2026Haowei Lou, Junda Wu, Chengkai Huang +4TTS SynthesisRepresentation Disentanglement

  5. WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

    Jul 7, 2026Sihang Nie, Jinxin Ji, Xiaofen Xing +4TTS SynthesisSpeech Prosody

  6. GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

    Jul 2, 2026Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4TTS SynthesisZero-Shot TTS

  7. A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

    Jul 1, 2026Siyi Wang, James Bailey, Ting DangLanguage Model SteeringConditional Flow Matching

  8. UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

    Jun 30, 2026Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong +4Audio EditingSpeech Editing

  9. SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

    Jun 29, 2026Zhuhan Bao, Rui Yang, Bohao Yang +16Human Behavior SimulationControllable Speech Generation

  10. VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

    Jun 25, 2026Tianxin Xie, Chenxing Li, Dong Yu +1Test-Time AdaptationReinforcement Learning

  11. Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

    Jun 22, 2026Seymanur Akti, Alexander WaibelTTS SynthesisSpeech Prosody

  12. Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

    Jun 22, 2026Jinchuan Tian, Haoran Wang, Siddhant Arora +6Controllable Speech GenerationText-to-Speech

  13. ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech

    Jun 20, 2026Wei Xue, Junlan Feng, Shilei Zhang +9TTS SynthesisTTS Evaluation

  14. Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

    Jun 17, 2026Michael Finkelson, Daniel Segal, Eitan Richardson +7TTS SynthesisControllable Speech Generation

  15. FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

    Jun 17, 2026Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4TTS SynthesisControllable Speech Generation

  16. DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention

    Jun 16, 2026Wenqiu Tang, Zhen Wan, Takahiro Komamizu +1Large Language Model-Based Role-Play SimulationSpoken Dialogue Systems

  17. Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

    Jun 11, 2026Yihang Lin, Li Zhou, Congwei Cao +4TTS SynthesisPreference Optimization

  18. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    Jun 8, 2026Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1TTS SynthesisLanguage Model Steering

  19. EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    Jun 8, 2026Minghui Wu, Ganjun Liu, Zikun Fang +6Representation LearningTTS Synthesis

  20. TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion

    Jun 5, 2026Constantin Alexander AugaLatent Diffusion ModelsAudio Diffusion Models

  21. VoxCPM2 Technical Report

    Jun 5, 2026Yixuan Zhou, Guoyang Zeng, Xin Liu +15Speech Foundation ModelsControllable Speech Generation

  22. GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

    Jun 4, 2026Jaehoon Kang, Yejin Lee, Kyuhong ShimTTS SynthesisAdapter Tuning

  23. Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech

    Jun 3, 2026Daniel Oliveira de Brito, Arnaldo Candido JuniorTTS SynthesisTask Arithmetic