TTS Synthesis

TTS: Text-to-Speech

Momentum

41 papers in the last four weeks, up 242% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 178

All topics
CardsList
  1. ZONOS2 Technical Report

    Jun 23, 2026Gabriel Clark, Sofian Mejjoute, Mohamed Osman +2TTS SynthesisStreaming TTS Synthesis

  2. Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

    Jun 22, 2026Seymanur Akti, Alexander WaibelTTS SynthesisSpeech Prosody

  3. ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech

    Jun 20, 2026Wei Xue, Junlan Feng, Shilei Zhang +9TTS SynthesisTTS Evaluation

  4. Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

    Jun 20, 2026Muyang Du, Jason Roche, Junjie LaiTTS SynthesisStreaming TTS Synthesis

  5. Sexualised synthetic personas encode and amplify gendered power asymmetries through voice

    Jun 19, 2026Alice Ross, Ariadna Sanchez, Elin Kanhov +2TTS SynthesisAlgorithmic Bias

  6. Imitation Learning for Elder-Facing Speech Synthesis

    Jun 19, 2026Dongrui Han, Weidong Chen, Jiawen Kang +3Reinforcement LearningTTS Synthesis

  7. FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

    Jun 18, 2026Harshit Singh, Ayush Pratap Singh, Nityanand MathurTTS SynthesisContinual Learning

  8. ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

    Jun 18, 2026Maxim Melichov, Yakov Kolani, Morris AlperTTS SynthesisLow-Resource Language Processing

  9. Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

    Jun 18, 2026Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1TTS SynthesisLow-Resource TTS Synthesis

  10. Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

    Jun 17, 2026Michael Finkelson, Daniel Segal, Eitan Richardson +7TTS SynthesisControllable Speech Generation

  11. FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

    Jun 17, 2026Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4TTS SynthesisControllable Speech Generation

  12. MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

    Jun 16, 2026Subhankar Ghosh, Jason Li, Paarth Neekhara +4TTS SynthesisSpeech Prosody

  13. Joycent: Multi-Accent TTS via Disentangled Accent Modeling and Layer-Specific Conditioning

    Jun 15, 2026Xintong Wang, Junchuan Zhao, Ye WangDisentangled Representation LearningTTS Synthesis

  14. Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

    Jun 14, 2026Yue Heng Yeo, Haoyang Li, Yizhou Peng +6Code-Switching Speech RecognitionSynthetic Data Augmentation

  15. Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

    Jun 13, 2026Zhenwei Mou, Liping Chen, Yajun Hu +3TTS SynthesisSpeech Prosody

  16. DuraMark: Duration-Embedded Watermarking in LLM-based TTS

    Jun 13, 2026Zhenwei Mou, Weili Jiang, Liping Chen +4Audio WatermarkingTTS Synthesis

  17. Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech

    Jun 12, 2026Alef Iury Siqueira Ferreira, Lucas Rafael Stefanel Gris, Luiz Fernando de Araújo Vidal +4Flow MatchingTTS Synthesis

  18. Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

    Jun 11, 2026Yihang Lin, Li Zhou, Congwei Cao +4TTS SynthesisPreference Optimization

  19. UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

    Jun 10, 2026Sangmin Lee, Eekgyun Ahn, Woongjib Choi +1TTS Synthesis

  20. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    Jun 8, 2026Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1TTS SynthesisLanguage Model Steering

  21. Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech

    Jun 8, 2026Vadim Popov, Wenju Gu, Tasnima Sadekova +2TTS SynthesisDiffusion Models

  22. OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

    Jun 8, 2026David Guzmán, Luel Hagos Beyene, Jesujoba Oluwadara Alabi +3TTS SynthesisTTS Evaluation