Zero-Shot TTS

TTS: Text-to-Speech

Momentum

10 papers in the last four weeks, up 233% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 42

All topics
CardsList
  1. Beyond Token Revision: Investigating Mask-and-Replace Diffusion for Zero-Shot Text-to-Speech

    Oct 7, 2026Hounsu Kim, Joonyong Park, Yuki Saito +2Zero-Shot TTSAudio Diffusion Models

  2. RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning

    Sep 29, 2026Maxim Maslov, Kirill Borodin, Vasilii Kudryavtsev +2Audio Diffusion ModelsZero-Shot TTS

  3. DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS

    Sep 26, 2026Ambuj Mehrish, Abhinaba Roy, Alex Ivanov +2TTS SynthesisZero-Shot TTS

  4. EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows

    Sep 24, 2026Hongyao Deng, Wenhao Guan, Xuetao Lin +4TTS SynthesisSpeech Editing

  5. Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech

    Sep 23, 2026Hyoeun Kim, Yujun Lee, Kyuhong ShimZero-Shot TTSMachine Unlearning

  6. TTS-Guard: Black-Box Ownership Verification of Text-to-Speech Models via Adaptive Adversarial Speaker-Pair Fingerprints

    Sep 20, 2026Xubin Yue, Zhenhua Xu, Zhebo Wang +5Zero-Shot TTS

  7. Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning

    Sep 14, 2026Qingyu Liu, Rixi Xu, Yushen Chen +11TTS SynthesisZero-Shot TTS

  8. StepAudio 3 Gen Technical Report

    Sep 14, 2026Bin Lin, Bo Zhao, Boyang Wang +68TTS SynthesisZero-Shot TTS

  9. Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

    Sep 10, 2026Lianru Gao, Yujie Guo, Yong QinTTS SynthesisZero-Shot TTS

  10. Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

    Aug 12, 2026Huaxuan Wang, Huimin Wang, Ruiyu Zhang +2TTS SynthesisZero-Shot TTS

  11. Luna-TTS Family Technical Report

    Aug 12, 2026Feng Yin, Shuai Shi, Junjie Zheng +19Non-Autoregressive Text GenerationTTS Synthesis

  12. CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

    Aug 9, 2026Yuqian Zhang, Yao Shi, Kexin Huang +6TTS SynthesisStreaming TTS Synthesis

  13. CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

    Aug 8, 2026Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5TTS SynthesisZero-Shot TTS

  14. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

    Jul 29, 2026Carlos Muñoz-Romero, Jose A. Gonzalez-LopezTTS SynthesisZero-Shot TTS

  15. GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

    Jul 2, 2026Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4TTS SynthesisZero-Shot TTS

  16. VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

    Jun 25, 2026Tianxin Xie, Chenxing Li, Dong Yu +1Test-Time AdaptationReinforcement Learning

  17. Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

    Jun 24, 2026Runwu Shi, Yujin Wang, Hongjin Song +8Flow MatchingClassifier-Free Guidance

  18. An Evaluation Framework for Text-to-Speech Voice Reconstruction

    Jun 19, 2026Ariadna Sanchez, Christoph Minixhofer, Korin Richmond +3Speech Generation EvaluationTTS Evaluation

  19. SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

    Jun 10, 2026Peijie Chen, Wenhao Guan, Weijie Wu +7Variational AutoencodersAudio Representation Learning

  20. BareWave: Waveform-Native Flow-Matching Text-to-Speech

    Jun 8, 2026Wei Fan, Chao-Hong Tan, Qian Chen +5Flow MatchingTTS Synthesis

  21. Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

    Jun 5, 2026Adarsh Arigala, Arjun Gangwar, S Umesh +1TTS SynthesisZero-Shot TTS

  22. GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

    Jun 4, 2026Jaehoon Kang, Yejin Lee, Kyuhong ShimTTS SynthesisAdapter Tuning

  23. WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

    Jun 2, 2026Wenxi Chen, Dongya Jia, Yushen Chen +11TTS SynthesisAudio Diffusion Models

  24. Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

    May 29, 2026Deokjin Seo, Gangin Park, Kihyun NamTTS SynthesisStreaming TTS Synthesis

  25. PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

    May 26, 2026Bowen Li, Shaotong Guo, Zhen Wang +11TTS SynthesisZero-Shot TTS

  26. Continual Speaker Identity Unlearning with Minimal Interference

    May 25, 2026Jinju Kim, Yunsung Kang, Gyeong-Moon Park +1TTS SynthesisZero-Shot TTS

  27. CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

    May 25, 2026Junyang Chen, Yuhang Jia, Hui Wang +3Speech EditingZero-Shot TTS

  28. RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

    May 21, 2026Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3Flow MatchingTTS Synthesis

  29. Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

    May 10, 2026Dong Yang, Yiyi Cai, Haoyu Zhang +2Flow MatchingTTS Synthesis

  30. X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

    May 7, 2026Rixi Xu, Qingyu Liu, Haitao Li +10TTS SynthesisZero-Shot TTS

  31. Speech Generation Speaker Poisoning: Capability Erasure in Zero-Shot Text-to-Speech

    Mar 8, 2026Thanathai Lertpetchpun, Thanapat Trachu, Sai Praneeth Karimireddy +1TTS SynthesisZero-Shot TTS

  32. Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

    Jan 19, 2026Seymanur Akti, Alexander WaibelTTS SynthesisZero-Shot TTS