TTS Synthesis

TTS: Text-to-Speech

Momentum

41 papers in the last four weeks, up 242% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 178

All topics
CardsList
  1. SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

    May 1, 2026Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2Representation LearningTTS Synthesis

  2. JaiTTS: A Thai Voice Cloning Model

    Apr 30, 2026Jullajak Karnjanaekarin, Pontakorn Trakuekul, Narongkorn Panitsrisit +5TTS SynthesisAudio Generation

  3. The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

    Apr 29, 2026Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou +5TTS SynthesisSpeech Generation Evaluation

  4. One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

    Apr 28, 2026Amanuel Gizachew Abebe, Yasmin MoslemTTS SynthesisText-to-Speech

  5. V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data

    Apr 25, 2026Tanusree Sharma, Anish Krishnagiri, Lili Dudas +2TTS SynthesisAudio Generation

  6. TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

    Apr 24, 2026Xi Wang, Jie Wang, Xingchen Song +8TTS SynthesisSpeech Generation Evaluation

  7. UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

    Apr 24, 2026Chunyu Qiang, Xiaopeng Wang, Kang Yin +11Text-to-Music GenerationText-to-Audio Generation

  8. Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

    Apr 23, 2026Srija Anand, Ashwin Sankar, Ishvinder Sethi +10Pairwise Preference LearningHuman Preference Evaluation

  9. MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

    Apr 23, 2026Jialong Mai, Xiaofen Xing, Xiangmin XuTTS SynthesisSpeech Prosody

  10. Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment

    Apr 21, 2026Huanchen Cai, Sten TernströmTTS SynthesisSpeech Quality Assessment

  11. ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

    Apr 21, 2026Aoduo Li, Haoran Lv, Hongjian Xu +5TTS SynthesisControllable Speech Generation

  12. MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

    Apr 20, 2026Huakang Chen, Jingbin Hu, Liumeng Xue +12TTS SynthesisSpeech Generation Evaluation

  13. AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

    Apr 17, 2026Sihan Lv, Yechen Jin, Zhen Li +5TTS SynthesisAudio Editing

  14. Acoustic and perceptual differences between standard and accented speech and their voice clones

    Apr 2, 2026Tianle Yang, Chengzhe Sun, Phil Rose +1TTS SynthesisAudio Generation

  15. SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation

    Mar 23, 2026Lucas H. Ueda, João G. T. Lima, Pedro R. Corrêa +3Disentangled Representation LearningTTS Synthesis

  16. WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    Mar 17, 2026Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1TTS SynthesisSelf-Attention

  17. Speech Generation Speaker Poisoning: Capability Erasure in Zero-Shot Text-to-Speech

    Mar 8, 2026Thanathai Lertpetchpun, Thanapat Trachu, Sai Praneeth Karimireddy +1TTS SynthesisZero-Shot TTS

  18. CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

    Feb 23, 2026Hanwen Liu, Saierdaer Yusuyin, Hao Huang +1TTS SynthesisStreaming TTS Synthesis

  19. Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

    Jan 19, 2026Seymanur Akti, Alexander WaibelTTS SynthesisZero-Shot TTS

  20. Decoding Order Matters in Autoregressive Speech Synthesis

    Jan 13, 2026Minghui Zhao, Anton RagniAdaptive InferenceTTS Synthesis

  21. SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

    Nov 23, 2025Kaidi Wang, Yi He, Wenhao Guan +9TTS SynthesisAudio-Visual Synchronization

  22. ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

    Oct 12, 2025Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh ShakeryTTS SynthesisSpeech Processing

  23. Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

    Jun 14, 2025Yakov Kolani, Maxim Melichov, Cobi Calev +1TTS SynthesisG2P Conversion