Controllable Speech Generation

Momentum

21 papers in the last four weeks, up 600% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 84

All topics
CardsList
  1. Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

    May 31, 2026Hongfei Du, Jiacheng Shi, Sidi Lu +2TTS SynthesisSparse Autoencoders

  2. Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

    May 30, 2026Sukru Samet Dindar, Riki Shimizu, Xilin Jiang +1Spoken Dialogue SystemsControllable Speech Generation

  3. HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

    May 28, 2026Bohan Li, Shi Lian, Hankun Wang +6Audio Representation LearningSpeech Foundation Models

  4. PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

    May 26, 2026Bowen Li, Shaotong Guo, Zhen Wang +11TTS SynthesisZero-Shot TTS

  5. Can We Hear from Events? Generating Speech from Event Camera

    May 26, 2026Jingping Fang, Lin Chen, Chenyang Xu +3Multimodal GenerationControllable Speech Generation

  6. DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

    May 20, 2026Xu Zhang, Longbing Cao, Zhangkai WuTTS SynthesisAudio Diffusion Models

  7. Bridging the Gap: Converting Read Text to Conversational Dialogue

    May 18, 2026Parshav Singla, Agnik Banerjee, Aaditya Arora +5TTS SynthesisSpeech Prosody

  8. Voice "Cloning" is Style Transfer

    May 15, 2026Kaitlyn Zhou, Federico Bianchi, Martijn Bartelds +3Controllable Speech GenerationTrust in AI

  9. AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

    May 14, 2026Bin Kang, Shaoguo Wen, Yang Fan +6TTS SynthesisControllable Speech Generation

  10. SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

    May 1, 2026Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2Representation LearningTTS Synthesis

  11. EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

    Apr 29, 2026Shuhao Xu, Yifan Hu, Jingjing Wu +3Emotion Recognition in ConversationsControllable Speech Generation

  12. MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

    Apr 23, 2026Jialong Mai, Xiaofen Xing, Xiangmin XuTTS SynthesisSpeech Prosody

  13. ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

    Apr 21, 2026Aoduo Li, Haoran Lv, Hongjian Xu +5TTS SynthesisControllable Speech Generation

  14. MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

    Apr 20, 2026Huakang Chen, Jingbin Hu, Liumeng Xue +12TTS SynthesisSpeech Generation Evaluation

  15. NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation

    Apr 17, 2026Liumeng Xue, Weizhen Bian, Jiahao Pan +9Speech Generation EvaluationSpeech Quality Assessment

  16. AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

    Apr 17, 2026Sihan Lv, Yechen Jin, Zhen Li +5TTS SynthesisAudio Editing

  17. OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

    Mar 25, 2026Seunghee Kim, Bumkyu Park, Kyudan Jung +5Multimodal GroundingUnified Multimodal Models

  18. SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation

    Mar 23, 2026Lucas H. Ueda, João G. T. Lima, Pedro R. Corrêa +3Disentangled Representation LearningTTS Synthesis

  19. Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

    Jan 19, 2026Seymanur Akti, Alexander WaibelTTS SynthesisZero-Shot TTS

  20. CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

    Jan 8, 2026Junyang Chen, Yuhang Jia, Hui Wang +2Audio EditingSpeech Editing

  21. PitchFlower: A flow-based neural audio codec with pitch controllability

    Oct 29, 2025Diego Torres, Axel Roebel, Nicolas ObinNeural Audio CodecsControllable Speech Generation