TTS Synthesis

TTS: Text-to-Speech

Momentum

41 papers in the last four weeks, up 242% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 178

All topics
CardsList
  1. NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech

    Jun 8, 2026Hongkun Yang, Xinhui Yi, Xiyan Zhao +13TTS SynthesisLow-Resource TTS Synthesis

  2. End-to-End Training for Discrete Token LLM based TTS System

    Jun 8, 2026Changfeng Gao, Yong Ren, Jun Yuan +3TTS SynthesisSpeech Processing

  3. EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    Jun 8, 2026Minghui Wu, Ganjun Liu, Zikun Fang +6Representation LearningTTS Synthesis

  4. BareWave: Waveform-Native Flow-Matching Text-to-Speech

    Jun 8, 2026Wei Fan, Chao-Hong Tan, Qian Chen +5Flow MatchingTTS Synthesis

  5. KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026

    Jun 5, 2026Seymanur Akti, Alexander WaibelTTS SynthesisCross-Lingual Transfer

  6. Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

    Jun 5, 2026Adarsh Arigala, Arjun Gangwar, S Umesh +1TTS SynthesisZero-Shot TTS

  7. dots.tts Technical Report

    Jun 5, 2026Shi Lian, Changtao Li, Bohan Li +6TTS SynthesisSpeech Foundation Models

  8. Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

    Jun 4, 2026Naman Kothari, Arjun Gangwar, Adarsh Arigala +1TTS SynthesisSpeech Processing

  9. GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

    Jun 4, 2026Jaehoon Kang, Yejin Lee, Kyuhong ShimTTS SynthesisAdapter Tuning

  10. UniVoice: A Unified Model for Speech and Singing Voice Generation

    Jun 4, 2026Junjie Zheng, Huixin Xue, Shihong Ren +3Singing Voice SynthesisTTS Synthesis

  11. Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech

    Jun 3, 2026Daniel Oliveira de Brito, Arnaldo Candido JuniorTTS SynthesisTask Arithmetic

  12. WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

    Jun 2, 2026Wenxi Chen, Dongya Jia, Yushen Chen +11TTS SynthesisAudio Diffusion Models

  13. UniVocal: Unified Speech-Singing Code-Switching Synthesis

    Jun 1, 2026Yufei Shi, Qian Chen, Wen Wang +3Singing Voice SynthesisTTS Synthesis

  14. Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

    May 31, 2026Hongfei Du, Jiacheng Shi, Sidi Lu +2TTS SynthesisSparse Autoencoders

  15. UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

    May 29, 2026Zhaoqing Li, Haoning Xu, Jingran Su +9Text-to-Audio GenerationTTS Synthesis

  16. ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

    May 29, 2026Jun-Hak Yun, Seung-Bin Kim, Seong-Whan LeeCross-Modal Representation LearningTTS Synthesis

  17. Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

    May 29, 2026Deokjin Seo, Gangin Park, Kihyun NamTTS SynthesisStreaming TTS Synthesis

  18. PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

    May 26, 2026Bowen Li, Shaotong Guo, Zhen Wang +11TTS SynthesisZero-Shot TTS

  19. Continual Speaker Identity Unlearning with Minimal Interference

    May 25, 2026Jinju Kim, Yunsung Kang, Gyeong-Moon Park +1TTS SynthesisZero-Shot TTS

  20. RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

    May 21, 2026Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3Flow MatchingTTS Synthesis

  21. Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

    May 20, 2026Vinicius Ribeiro, Yves LaprieTTS SynthesisSpeech Quality Assessment

  22. DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

    May 20, 2026Xu Zhang, Longbing Cao, Zhangkai WuTTS SynthesisAudio Diffusion Models

  23. Bridging the Gap: Converting Read Text to Conversational Dialogue

    May 18, 2026Parshav Singla, Agnik Banerjee, Aaditya Arora +5TTS SynthesisSpeech Prosody

  24. AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

    May 14, 2026Bin Kang, Shaoguo Wen, Yang Fan +6TTS SynthesisControllable Speech Generation

  25. AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling

    May 12, 2026Yiming Ren, Xuenan Xu, Ziyang Zhang +3TTS SynthesisHuman-in-the-Loop AI

  26. Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

    May 10, 2026Dong Yang, Yiyi Cai, Haoyu Zhang +2Flow MatchingTTS Synthesis

  27. WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

    May 7, 2026Guanrou Yang, Tian Tan, Qian Chen +12Audio Representation LearningTTS Synthesis

  28. X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

    May 7, 2026Rixi Xu, Qingyu Liu, Haitao Li +10TTS SynthesisZero-Shot TTS

  29. Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

    May 4, 2026Jiaxu He, Chao Wang, Jie Lian +4TTS SynthesisLow-Resource TTS Synthesis