TTS Synthesis

TTS: Text-to-Speech

Momentum

41 papers in the last four weeks, up 242% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 178

All topics
CardsList
  1. Steerspeech: Activation Steering For Emotion Control In Generated Speech

    Oct 7, 2026Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1Controllable Speech GenerationEmotional Speech Synthesis

  2. Training-Free Instruction TTS Gender Bias Calibration Using Model-Adaptive Steering

    Oct 7, 2026Kuan-Yu Chen, Yi-Cheng Lin, Jeng-Lin Li +1Text-to-SpeechTTS Synthesis

  3. Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

    Oct 6, 2026Seymanur Akti, Alexander WaibelTTS SynthesisControllable Speech Generation

  4. Pronunciation-Oriented Reinforcement Learning for Japanese Text-to-Speech with Kana-Domain ASR Rewards

    Oct 6, 2026Shiao Zhu, Lianbo Liu, Kai Washizaki +2TTS Synthesis

  5. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

    Oct 4, 2026Abdul Rehman, Jian-Jun Zhang, Xiaosong YangTTS SynthesisSpeech Processing

  6. Balalaika-Longform: A Russian Speech Corpus for Continuous Long-Form Text-to-Speech

    Sep 30, 2026Nikita Vasiliev, Kirill Borodin, Vasilii Kudryavtsev +2TTS SynthesisTTS Evaluation

  7. SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS

    Sep 30, 2026Longyu Lu, Zongwei Du, Mengtao Xing +4TTS SynthesisStreaming TTS Synthesis

  8. RVQ Position Aware Speculative Decoding for On Device Text to Speech

    Sep 29, 2026Berkin Durmus, Eduardo Pacheco, Zach Nagengast +1TTS SynthesisStreaming TTS Synthesis

  9. Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech

    Sep 29, 2026Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov +1TTS SynthesisTTS Evaluation

  10. Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

    Sep 28, 2026Yentl Collin, Evan Dufraisse, Amr Mohamed +3Flow MatchingTTS Synthesis

  11. InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech

    Sep 28, 2026Sihang Nie, Xueru Li, Xiaofen Xing +4TTS SynthesisLanguage Model-Based Control

  12. SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows

    Sep 28, 2026Tianxin Xie, Pengfei Zhang, Kai Jiang +2TTS SynthesisSpeech Editing

  13. Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

    Sep 28, 2026Isha Pandey, Varad Deshpande, Abhijat Bharadwaj +1TTS SynthesisConditional Flow Matching

  14. Controlling Speaking Rate in Autoregressive TTS via Activation Steering

    Sep 27, 2026Francesco Verdini, Antonis Asonitis, Aref Farhadipour +4TTS SynthesisControllable Speech Generation

  15. DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS

    Sep 26, 2026Ambuj Mehrish, Abhinaba Roy, Alex Ivanov +2TTS SynthesisZero-Shot TTS

  16. EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows

    Sep 24, 2026Hongyao Deng, Wenhao Guan, Xuetao Lin +4TTS SynthesisSpeech Editing

  17. ReaFlow-TTS: Realization-Conditioned Flow Matching for High-Quality and Controllable Speech Synthesis

    Sep 24, 2026Junyi Zhao, Yihao Qin, Changsheng MaFlow MatchingTTS Synthesis

  18. EmphTTS: an emphasis-control TTS with reinforcement learning

    Sep 23, 2026Zirui Li, Rech Silas, Lauri Juvela +2TTS SynthesisControllable Speech Generation

  19. Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

    Sep 22, 2026Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova +3TTS SynthesisAudio-Visual Synchronization

  20. CycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding

    Sep 21, 2026Huan Liao, Haonan Han, Xingwen Han +3TTS SynthesisControllable Speech Generation

  21. Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech

    Sep 21, 2026Mizbaul Haque Maruf, Muhammad Nur YanhaonaTTS SynthesisCoreset Selection

  22. Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis

    Sep 21, 2026Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4TTS SynthesisAudio Reasoning

  23. HaikuS2S: A Cascaded System For Responding In Verse

    Sep 20, 2026Devangi Sharma, Sophia Judicke, Glenda Tan +2TTS SynthesisSpeech Generation Evaluation

  24. GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

    Sep 16, 2026Zitao Liang, Chang GaoTTS SynthesisSpeech Processing

  25. Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech

    Sep 15, 2026Shuhei KatoTTS SynthesisSpeech Prosody

  26. Taming Long-form Text-to-Speech

    Sep 15, 2026Rongxiang Wang, Berkin Durmus, Aysegul Orhon +2TTS SynthesisTTS Evaluation

  27. Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning

    Sep 14, 2026Qingyu Liu, Rixi Xu, Yushen Chen +11TTS SynthesisZero-Shot TTS

  28. StepAudio 3 Gen Technical Report

    Sep 14, 2026Bin Lin, Bo Zhao, Boyang Wang +68TTS SynthesisZero-Shot TTS

  29. Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech

    Sep 13, 2026Moses Daudu, Adeola Enitan Bamidele, Honor-Jesus BezaleelTTS SynthesisTTS Evaluation

  30. Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

    Sep 12, 2026Mattias Cross, Minghui Zhao, Anton RagniNeural Controlled Differential EquationsTTS Synthesis

  31. Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

    Sep 11, 2026Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2TTS SynthesisTTS Evaluation

  32. Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

    Sep 10, 2026Lianru Gao, Yujie Guo, Yong QinTTS SynthesisZero-Shot TTS

  33. Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

    Sep 9, 2026Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis +1TTS SynthesisSpeech Processing

  34. TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction

    Sep 8, 2026Yi-Chang Chen, Chun Wei Chen, Dien-Ruei Wu +7TTS SynthesisStreaming TTS Synthesis

  35. KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

    Sep 7, 2026Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji OgawaTTS SynthesisControllable Speech Generation

  36. Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

    Sep 3, 2026Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong +3TTS SynthesisTTS Evaluation

  37. Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

    Sep 1, 2026Kunlin Cai, Kaiyuan Zhang, Zihang Xiang +4TTS SynthesisMembership Inference Attacks

  38. Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech

    Sep 1, 2026Che Hyun Lee, Sangkwon Park, Donghun Kang +4TTS Synthesis

  39. Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation

    Aug 27, 2026Zhen Wang, TianRui Wu, RongQi Han +3Synthetic Data AugmentationTTS Synthesis

  40. Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

    Aug 12, 2026Huaxuan Wang, Huimin Wang, Ruiyu Zhang +2TTS SynthesisZero-Shot TTS

  41. Luna-TTS Family Technical Report

    Aug 12, 2026Feng Yin, Shuai Shi, Junjie Zheng +19Non-Autoregressive Text GenerationTTS Synthesis

  42. CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

    Aug 12, 2026Haowei Lou, Hye-Young Paik, Dai Jia +2Singing Voice SynthesisTTS Synthesis

  43. CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

    Aug 9, 2026Yuqian Zhang, Yao Shi, Kexin Huang +6TTS SynthesisStreaming TTS Synthesis