Speech Prosody

Momentum

12 papers in the last four weeks, up 300% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 53

All topics
CardsList
  1. A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization

    Oct 6, 2026Tien-Hong Lo, Fong-Chun Tsai, Ting-An Hung +3Speech ProcessingSpeech Prosody

  2. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

    Oct 4, 2026Abdul Rehman, Jian-Jun Zhang, Xiaosong YangTTS SynthesisSpeech Processing

  3. A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning

    Sep 30, 2026Feng Xu, Gaoyuan Zhang, Shanshan Xue +4Speech ProsodySpeech Emotion Recognition

  4. EmphTTS: an emphasis-control TTS with reinforcement learning

    Sep 23, 2026Zirui Li, Rech Silas, Lauri Juvela +2TTS SynthesisControllable Speech Generation

  5. Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

    Sep 21, 2026João Lima, Lucas Ueda, Paula CostaSpeech ProcessingSpeech Prosody

  6. HaikuS2S: A Cascaded System For Responding In Verse

    Sep 20, 2026Devangi Sharma, Sophia Judicke, Glenda Tan +2TTS SynthesisSpeech Generation Evaluation

  7. Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

    Sep 17, 2026Zifan Guan, Longyu Lu, Junan Zhang +3Speech Generation EvaluationTTS Evaluation

  8. T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

    Sep 16, 2026Hung-Yang Sung, Chien-Chun Wang, Tien-Hong Lo +3Automatic Speech RecognitionSpeech Prosody

  9. Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech

    Sep 15, 2026Shuhei KatoTTS SynthesisSpeech Prosody

  10. CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection

    Sep 15, 2026Qiyang Sun, Xudong Li, Yupei Li +4Audio-Language Model EvaluationCounterfactual Evaluation

  11. Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech

    Sep 13, 2026Moses Daudu, Adeola Enitan Bamidele, Honor-Jesus BezaleelTTS SynthesisTTS Evaluation

  12. Less can be More: What Aspects of Speech Drive End-of-Turn Detection

    Sep 12, 2026Rini Sharon, Manickavela A, Kadri Hacioglu +1Speech ProsodyTurn-Taking Prediction

  13. LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

    Sep 9, 2026Xiao Wei, Yuqin Lin, Yaru Cao +6Speech-Based Dementia DetectionSpeech Prosody

  14. Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

    Sep 1, 2026Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh +1Audio-Language Model EvaluationAudio-Language Models

  15. Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments

    Aug 31, 2026Ivan Porupski, Nikola LjubešićPolitical Discourse AnalysisSpeech Prosody

  16. Using Prosody to Predict Syntactic Structure

    Aug 31, 2026Junghyun Min, Alex Warstadt, Tamar I. Regev +2Speech Prosody

  17. CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

    Aug 8, 2026Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5TTS SynthesisZero-Shot TTS

  18. Correlation between prosody and pragmatics: A case study of the discourse marker hālā `now' in Persian

    Jul 30, 2026Soleiman Ghaderi, Moloud Asakereh, Kevin TangSpeech Prosody

  19. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

    Jul 22, 2026Kaicheng Luo, Xuefei Gong, Yutao Sun +6TTS SynthesisSpeech Prosody

  20. Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

    Jul 15, 2026Stephen McIntosh, Reuben Smit, Daisuke Saito +2Dynamic Time WarpingSpeech Processing

  21. Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese

    Jul 8, 2026Rodrigo de Freitas Lima, Julio Cesar Galdino, Marcos Vinicius TrevisoSpeech ProsodySpeech Language Models

  22. WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

    Jul 7, 2026Sihang Nie, Jinxin Ji, Xiaofen Xing +4TTS SynthesisSpeech Prosody

  23. Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words

    Jul 2, 2026Xiaoyun Jin, Mirjam Ernestus, R. Harald BaayenSpeech ProcessingSpeech Prosody

  24. Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

    Jun 30, 2026Ashish Hallur, Thomas Thebaud, Georgi Tinchev +2Voice Agent EvaluationSpeech Generation Evaluation

  25. HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

    Jun 26, 2026Sihang Nie, Xiaofen Xing, Rui Xing +5Reward ModelingSpeech Prosody

  26. Do Speech Emphasis Models Generalize across Languages and Emotions?

    Jun 26, 2026Megan Wei, Deepali Aneja, Jiaqi Su +3Speech ProsodyCross-Lingual Transfer

  27. Phonetic and semantic analyses of spoken corpora of Beijing and Taiwan Mandarin indicate that the neutral tone is a lexical tone

    Jun 24, 2026Yuxin Lu, Zhexuan Li, R. Harald BaayenSpeech Prosody

  28. Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS

    Jun 24, 2026Sandipan Dhar, Nirmesh J. Shah, Ashishkumar P. Gudmalwar +1TTS SynthesisAudio Diffusion Models

  29. Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

    Jun 22, 2026Seymanur Akti, Alexander WaibelTTS SynthesisSpeech Prosody

  30. ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

    Jun 20, 2026Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu +1Audio Representation LearningNeural Audio Codecs

  31. LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations

    Jun 19, 2026Younghan Park, Hoyeon Lee, Hawon Jeong +1Speech ProsodyAutomated Evaluation

  32. MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

    Jun 16, 2026Subhankar Ghosh, Jason Li, Paarth Neekhara +4TTS SynthesisSpeech Prosody

  33. Perceptual compensation for tonal context in self-supervised speech models

    Jun 16, 2026James Kirby, Ioana Krehan, Michele GubianSupervised Fine-TuningSpeech Prosody

  34. Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

    Jun 13, 2026Zhenwei Mou, Liping Chen, Yajun Hu +3TTS SynthesisSpeech Prosody

  35. NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

    Jun 11, 2026Dongwook Lee, Youngho Cho, Sangkwon Park +2Speech TranslationSimultaneous Speech Translation

  36. PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

    Jun 11, 2026Wen Zhang, Xiaocui Yang, Zhuoyue Gao +3Multi-Agent LLM SystemsSpoken Dialogue Systems

  37. What Makes Synthetic Speech Sound Sarcastic? A Prosody-Controlled Perception Study

    Jun 8, 2026Zhu Li, Shekhar Nayak, Matt ColerSpeech Prosody

  38. ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

    Jun 4, 2026Prathamjyot Singh, Ashima Sood, Sahil Sharma +1Audio UnderstandingSpeech Prosody

  39. Task-Vector Arithmetic for Emotional Expressivity Control in Language-Model-Based Text-to-Speech

    Jun 3, 2026Daniel Oliveira de Brito, Arnaldo Candido JuniorTTS SynthesisTask Arithmetic

  40. Why We Need Speech to Evaluate Speech Translation

    May 27, 2026Maike Züfle, Danni Liu, Vilém Zouhar +1Speech TranslationSpeech Prosody

  41. Bridging the Gap: Converting Read Text to Conversational Dialogue

    May 18, 2026Parshav Singla, Agnik Banerjee, Aaditya Arora +5TTS SynthesisSpeech Prosody

  42. A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges

    May 11, 2026Wen Liang, Li Siyan, Zackary Rackauckas +1Speech Prosody

  43. Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

    May 7, 2026Wenqian Cui, Xiao-Hui Li, Daxin Tan +2Audio UnderstandingSpeech Prosody

  44. DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

    Apr 29, 2026Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews +2Audio Diffusion ModelsSpeech Prosody

  45. MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

    Apr 23, 2026Jialong Mai, Xiaofen Xing, Xiangmin XuTTS SynthesisSpeech Prosody

  46. Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean

    Apr 21, 2026Hyunjung Joo, GyeongTaek LeeSpeech ProsodySupervised Contrastive Learning

  47. AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design

    Apr 19, 2026Yuchen Liu, Roberto Martinez-Maldonado, Riordan Alfredo +3Speech Prosody

  48. Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations

    Apr 2, 2026Haitong Sun, Stephen McIntosh, Kwanghee Choi +3Speech ProsodySelf-Supervised Speech Representation Learning

  49. Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

    Jul 17, 2025Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev +3Speech ProcessingSpeech Prosody