TTS Synthesis

TTS: Text-to-Speech

Momentum

41 papers in the last four weeks, up 242% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 178

All topics
CardsList
  1. Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis

    Sep 21, 2026Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4TTS SynthesisAudio Reasoning

  2. HaikuS2S: A Cascaded System For Responding In Verse

    Sep 20, 2026Devangi Sharma, Sophia Judicke, Glenda Tan +2TTS SynthesisSpeech Generation Evaluation

  3. GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

    Sep 16, 2026Zitao Liang, Chang GaoTTS SynthesisSpeech Processing

  4. Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech

    Sep 15, 2026Shuhei KatoTTS SynthesisSpeech Prosody

  5. Taming Long-form Text-to-Speech

    Sep 15, 2026Rongxiang Wang, Berkin Durmus, Aysegul Orhon +2TTS SynthesisTTS Evaluation

  6. Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning

    Sep 14, 2026Qingyu Liu, Rixi Xu, Yushen Chen +11TTS SynthesisZero-Shot TTS

  7. StepAudio 3 Gen Technical Report

    Sep 14, 2026Bin Lin, Bo Zhao, Boyang Wang +68TTS SynthesisZero-Shot TTS

  8. Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech

    Sep 13, 2026Moses Daudu, Adeola Enitan Bamidele, Honor-Jesus BezaleelTTS SynthesisTTS Evaluation

  9. Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

    Sep 12, 2026Mattias Cross, Minghui Zhao, Anton RagniNeural Controlled Differential EquationsTTS Synthesis

  10. Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

    Sep 11, 2026Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2TTS SynthesisTTS Evaluation

  11. Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

    Sep 10, 2026Lianru Gao, Yujie Guo, Yong QinTTS SynthesisZero-Shot TTS

  12. Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

    Sep 9, 2026Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis +1TTS SynthesisSpeech Processing

  13. TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction

    Sep 8, 2026Yi-Chang Chen, Chun Wei Chen, Dien-Ruei Wu +7TTS SynthesisStreaming TTS Synthesis

  14. KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

    Sep 7, 2026Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji OgawaTTS SynthesisControllable Speech Generation

  15. Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

    Sep 3, 2026Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong +3TTS SynthesisTTS Evaluation

  16. Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

    Sep 1, 2026Kunlin Cai, Kaiyuan Zhang, Zihang Xiang +4TTS SynthesisMembership Inference Attacks

  17. Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech

    Sep 1, 2026Che Hyun Lee, Sangkwon Park, Donghun Kang +4TTS Synthesis

  18. Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation

    Aug 27, 2026Zhen Wang, TianRui Wu, RongQi Han +3Synthetic Data AugmentationTTS Synthesis

  19. Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

    Aug 12, 2026Huaxuan Wang, Huimin Wang, Ruiyu Zhang +2TTS SynthesisZero-Shot TTS

  20. Luna-TTS Family Technical Report

    Aug 12, 2026Feng Yin, Shuai Shi, Junjie Zheng +19Non-Autoregressive Text GenerationTTS Synthesis

  21. CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

    Aug 12, 2026Haowei Lou, Hye-Young Paik, Dai Jia +2Singing Voice SynthesisTTS Synthesis

  22. CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

    Aug 9, 2026Yuqian Zhang, Yao Shi, Kexin Huang +6TTS SynthesisStreaming TTS Synthesis