Controllable Speech Generation

Momentum

21 papers in the last four weeks, up 600% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 84

All topics
CardsList
  1. Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

    Oct 8, 2026Shiao Zhu, Lianbo Liu, Sizhen Lyu +3TTS SynthesisControllable Speech Generation

  2. Edit Who Speaks, Control How They Speak: Global Timbre Editing and Local Instruction Control for TTS

    Oct 8, 2026Junchuan Zhao, Chenglin Xu, Wei Zeng +3TTS SynthesisControllable Speech Generation

  3. Steerspeech: Activation Steering For Emotion Control In Generated Speech

    Oct 7, 2026Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1Controllable Speech GenerationEmotional Speech Synthesis

  4. Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

    Oct 6, 2026Seymanur Akti, Alexander WaibelTTS SynthesisControllable Speech Generation

  5. Tracing a Sparse Emotion-Control Circuit in LLM-Based Text-to-Speech

    Oct 4, 2026Hongfei Du, Jiacheng Shi, Yanfu Zhang +1Mechanistic InterpretabilityControllable Speech Generation

  6. SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS

    Sep 30, 2026Longyu Lu, Zongwei Du, Mengtao Xing +4TTS SynthesisStreaming TTS Synthesis

  7. EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

    Sep 29, 2026Kuan-Po Huang, Haohe Liu, Puyuan Peng +5Controllable Speech GenerationEmotional Speech Synthesis

  8. InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech

    Sep 28, 2026Sihang Nie, Xueru Li, Xiaofen Xing +4TTS SynthesisLanguage Model-Based Control

  9. SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows

    Sep 28, 2026Tianxin Xie, Pengfei Zhang, Kai Jiang +2TTS SynthesisSpeech Editing

  10. Controlling Speaking Rate in Autoregressive TTS via Activation Steering

    Sep 27, 2026Francesco Verdini, Antonis Asonitis, Aref Farhadipour +4TTS SynthesisControllable Speech Generation

  11. ReaFlow-TTS: Realization-Conditioned Flow Matching for High-Quality and Controllable Speech Synthesis

    Sep 24, 2026Junyi Zhao, Yihao Qin, Changsheng MaFlow MatchingTTS Synthesis

  12. EmphTTS: an emphasis-control TTS with reinforcement learning

    Sep 23, 2026Zirui Li, Rech Silas, Lauri Juvela +2TTS SynthesisControllable Speech Generation

  13. Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

    Sep 22, 2026Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova +3TTS SynthesisAudio-Visual Synchronization

  14. CycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding

    Sep 21, 2026Huan Liao, Haonan Han, Xingwen Han +3TTS SynthesisControllable Speech Generation

  15. StepAudio 3 Gen Technical Report

    Sep 14, 2026Bin Lin, Bo Zhao, Boyang Wang +68TTS SynthesisZero-Shot TTS

  16. Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

    Sep 10, 2026Lianru Gao, Yujie Guo, Yong QinTTS SynthesisZero-Shot TTS

  17. Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation

    Sep 10, 2026Jingbin Hu, Qirui Zhan, Yuang Cao +7Direct Preference OptimizationControllable Speech Generation

  18. TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction

    Sep 8, 2026Yi-Chang Chen, Chun Wei Chen, Dien-Ruei Wu +7TTS SynthesisStreaming TTS Synthesis

  19. Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering

    Sep 8, 2026Yizhong Geng, Kecan Mao, Qifei Li +6TTS EvaluationTraining Data Curation

  20. KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

    Sep 7, 2026Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji OgawaTTS SynthesisControllable Speech Generation

  21. Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS

    Aug 31, 2026Yan Zhou, Yun Hong, Yang FengControllable Speech GenerationEmotional Speech Synthesis

  22. CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

    Aug 12, 2026Haowei Lou, Hye-Young Paik, Dai Jia +2Singing Voice SynthesisTTS Synthesis

  23. CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

    Aug 8, 2026Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5TTS SynthesisZero-Shot TTS