TTS Synthesis

TTS: Text-to-Speech

Momentum

41 papers in the last four weeks, up 242% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 178

All topics
CardsList
  1. NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech

    Jun 8, 2026Hongkun Yang, Xinhui Yi, Xiyan Zhao +13TTS SynthesisLow-Resource TTS Synthesis

  2. End-to-End Training for Discrete Token LLM based TTS System

    Jun 8, 2026Changfeng Gao, Yong Ren, Jun Yuan +3TTS SynthesisSpeech Processing

  3. EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    Jun 8, 2026Minghui Wu, Ganjun Liu, Zikun Fang +6Representation LearningTTS Synthesis

  4. BareWave: Waveform-Native Flow-Matching Text-to-Speech

    Jun 8, 2026Wei Fan, Chao-Hong Tan, Qian Chen +5Flow MatchingTTS Synthesis

  5. KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026

    Jun 5, 2026Seymanur Akti, Alexander WaibelTTS SynthesisCross-Lingual Transfer

  6. Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

    Jun 5, 2026Adarsh Arigala, Arjun Gangwar, S Umesh +1TTS SynthesisZero-Shot TTS

  7. dots.tts Technical Report

    Jun 5, 2026Shi Lian, Changtao Li, Bohan Li +6TTS SynthesisSpeech Foundation Models

  8. Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

    Jun 4, 2026Naman Kothari, Arjun Gangwar, Adarsh Arigala +1TTS SynthesisSpeech Processing

  9. GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

    Jun 4, 2026Jaehoon Kang, Yejin Lee, Kyuhong ShimTTS SynthesisAdapter Tuning

  10. UniVoice: A Unified Model for Speech and Singing Voice Generation

    Jun 4, 2026Junjie Zheng, Huixin Xue, Shihong Ren +3Singing Voice SynthesisTTS Synthesis

  11. Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech

    Jun 3, 2026Daniel Oliveira de Brito, Arnaldo Candido JuniorTTS SynthesisTask Arithmetic

  12. WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

    Jun 2, 2026Wenxi Chen, Dongya Jia, Yushen Chen +11TTS SynthesisAudio Diffusion Models

  13. UniVocal: Unified Speech-Singing Code-Switching Synthesis

    Jun 1, 2026Yufei Shi, Qian Chen, Wen Wang +3Singing Voice SynthesisTTS Synthesis

  14. Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

    May 31, 2026Hongfei Du, Jiacheng Shi, Sidi Lu +2TTS SynthesisSparse Autoencoders

  15. UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

    May 29, 2026Zhaoqing Li, Haoning Xu, Jingran Su +9Text-to-Audio GenerationTTS Synthesis

  16. ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

    May 29, 2026Jun-Hak Yun, Seung-Bin Kim, Seong-Whan LeeCross-Modal Representation LearningTTS Synthesis

  17. Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

    May 29, 2026Deokjin Seo, Gangin Park, Kihyun NamTTS SynthesisStreaming TTS Synthesis

  18. PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

    May 26, 2026Bowen Li, Shaotong Guo, Zhen Wang +11TTS SynthesisZero-Shot TTS

  19. Continual Speaker Identity Unlearning with Minimal Interference

    May 25, 2026Jinju Kim, Yunsung Kang, Gyeong-Moon Park +1TTS SynthesisZero-Shot TTS

  20. RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

    May 21, 2026Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3Flow MatchingTTS Synthesis

  21. Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

    May 20, 2026Vinicius Ribeiro, Yves LaprieTTS SynthesisSpeech Quality Assessment

  22. DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

    May 20, 2026Xu Zhang, Longbing Cao, Zhangkai WuTTS SynthesisAudio Diffusion Models

  23. Bridging the Gap: Converting Read Text to Conversational Dialogue

    May 18, 2026Parshav Singla, Agnik Banerjee, Aaditya Arora +5TTS SynthesisSpeech Prosody

  24. AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

    May 14, 2026Bin Kang, Shaoguo Wen, Yang Fan +6TTS SynthesisControllable Speech Generation

  25. AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling

    May 12, 2026Yiming Ren, Xuenan Xu, Ziyang Zhang +3TTS SynthesisHuman-in-the-Loop AI

  26. Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

    May 10, 2026Dong Yang, Yiyi Cai, Haoyu Zhang +2Flow MatchingTTS Synthesis

  27. WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

    May 7, 2026Guanrou Yang, Tian Tan, Qian Chen +12Audio Representation LearningTTS Synthesis

  28. X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

    May 7, 2026Rixi Xu, Qingyu Liu, Haitao Li +10TTS SynthesisZero-Shot TTS

  29. Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

    May 4, 2026Jiaxu He, Chao Wang, Jie Lian +4TTS SynthesisLow-Resource TTS Synthesis

  30. SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

    May 1, 2026Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2Representation LearningTTS Synthesis

  31. JaiTTS: A Thai Voice Cloning Model

    Apr 30, 2026Jullajak Karnjanaekarin, Pontakorn Trakuekul, Narongkorn Panitsrisit +5TTS SynthesisAudio Generation

  32. The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

    Apr 29, 2026Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou +5TTS SynthesisSpeech Generation Evaluation

  33. One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

    Apr 28, 2026Amanuel Gizachew Abebe, Yasmin MoslemTTS SynthesisText-to-Speech

  34. V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data

    Apr 25, 2026Tanusree Sharma, Anish Krishnagiri, Lili Dudas +2TTS SynthesisAudio Generation

  35. TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

    Apr 24, 2026Xi Wang, Jie Wang, Xingchen Song +8TTS SynthesisSpeech Generation Evaluation

  36. UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

    Apr 24, 2026Chunyu Qiang, Xiaopeng Wang, Kang Yin +11Text-to-Music GenerationText-to-Audio Generation

  37. Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

    Apr 23, 2026Srija Anand, Ashwin Sankar, Ishvinder Sethi +10Pairwise Preference LearningHuman Preference Evaluation

  38. MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

    Apr 23, 2026Jialong Mai, Xiaofen Xing, Xiangmin XuTTS SynthesisSpeech Prosody

  39. Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment

    Apr 21, 2026Huanchen Cai, Sten TernströmTTS SynthesisSpeech Quality Assessment

  40. ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

    Apr 21, 2026Aoduo Li, Haoran Lv, Hongjian Xu +5TTS SynthesisControllable Speech Generation

  41. MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

    Apr 20, 2026Huakang Chen, Jingbin Hu, Liumeng Xue +12TTS SynthesisSpeech Generation Evaluation

  42. AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

    Apr 17, 2026Sihan Lv, Yechen Jin, Zhen Li +5TTS SynthesisAudio Editing

  43. Acoustic and perceptual differences between standard and accented speech and their voice clones

    Apr 2, 2026Tianle Yang, Chengzhe Sun, Phil Rose +1TTS SynthesisAudio Generation

  44. SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation

    Mar 23, 2026Lucas H. Ueda, João G. T. Lima, Pedro R. Corrêa +3Disentangled Representation LearningTTS Synthesis

  45. WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    Mar 17, 2026Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1TTS SynthesisSelf-Attention

  46. Speech Generation Speaker Poisoning: Capability Erasure in Zero-Shot Text-to-Speech

    Mar 8, 2026Thanathai Lertpetchpun, Thanapat Trachu, Sai Praneeth Karimireddy +1TTS SynthesisZero-Shot TTS

  47. CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

    Feb 23, 2026Hanwen Liu, Saierdaer Yusuyin, Hao Huang +1TTS SynthesisStreaming TTS Synthesis

  48. Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

    Jan 19, 2026Seymanur Akti, Alexander WaibelTTS SynthesisZero-Shot TTS

  49. Decoding Order Matters in Autoregressive Speech Synthesis

    Jan 13, 2026Minghui Zhao, Anton RagniAdaptive InferenceTTS Synthesis

  50. SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

    Nov 23, 2025Kaidi Wang, Yi He, Wenhao Guan +9TTS SynthesisAudio-Visual Synchronization

  51. ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

    Oct 12, 2025Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh ShakeryTTS SynthesisSpeech Processing

  52. Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

    Jun 14, 2025Yakov Kolani, Maxim Melichov, Cobi Calev +1TTS SynthesisG2P Conversion