Speech Processing

Latest papers 295

All topics
CardsList
  1. Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception

    May 21, 2026Nicolas M. Müller, Wei Herng ChoongAudio Deepfake DetectionSpeech Processing

  2. Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

    May 20, 2026Vinicius Ribeiro, Yves LaprieTTS SynthesisSpeech Quality Assessment

  3. Cross-Talk Speech Reduction, by Separation, for Separation

    May 19, 2026Zhong-Qiu Wang, Samuele CornellSpeech ProcessingSpeech Separation

  4. FormalASR: End-to-End Spoken Chinese to Formal Text

    May 19, 2026Wanyi Ning, Yinshang Guo, Haitao Qian +3Speech ProcessingAutomatic Speech Recognition

  5. PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions

    May 18, 2026Sicheng Jin, Dipankar Srirag, Aditya JoshiASR EvaluationSpeech Processing

  6. Speaker-Disentangled Remote Speech Detection of Asthma and COPD Exacerbations

    May 16, 2026Yuyang Yan, Sami O. Simons, Visara UroviSpeech ProcessingClinical Prediction

  7. Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

    May 15, 2026Si-Belkacem Yamine Ketir, Lenard Paulo Tamayo, Shohei Hisada +3Class-Imbalanced LearningSynthetic Data Augmentation

  8. Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues

    May 15, 2026Zhongjie Ba, Liang Yi, Peng Cheng +3Speech ProcessingHate Speech Detection

  9. SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

    May 14, 2026KiHyun Nam, Jungwoo Heo, Siu Bae +2Audio Representation LearningSpeaker Verification

  10. PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection

    May 14, 2026Madhurananda Pahar, Caitlin H. Illingworth, Bahman Mirheidari +7Speech-Based Dementia DetectionSpeech Processing

  11. A Benchmark for Early-stage Parkinson's Disease Detection from Speech

    May 13, 2026Terry Yi Zhong, Cristian Tejedor-Garcia, Khiet P. Truong +3Benchmark DesignSpeech Processing

  12. Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs

    May 12, 2026Deepak Kumar, Baban Gain, Asif EkbalSpeech Processing

  13. Predicting Psychological Well-Being from Spontaneous Speech using LLMs

    May 11, 2026Erfan Loweimi, Sofia de la Fuente Garcia, Saturnino LuzLLM EvaluationSpeech Processing

  14. Exploring Token-Space Manipulation in Latent Audio Tokenizers

    May 11, 2026Francesco Paissan, Luca Della Libera, Mirco Ravanelli +1Audio Representation LearningNeural Audio Codecs

  15. Voice Biomarkers for Depression and Anxiety

    May 11, 2026Oleksii Abramenko, Noah D. Stein, Colin VazAudio Representation LearningDepression Detection

  16. End-to-End Keyword Spotting on FPGA Using Graph Neural Networks with a Neuromorphic Auditory Sensor

    May 10, 2026Wiktor Matykiewicz, Piotr Wzorek, Kamil Jeziorek +4Keyword SpottingSpeech Processing

  17. Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

    May 8, 2026Michael Neri, Archontis Politis, Tuomas VirtanenSpeech ProcessingRoom Acoustics

  18. Linear Semantic Segmentation for Low-Resource Spoken Dialects

    May 7, 2026Kirill Chirkunov, Younes Samih, Abed Alhakim Freihat +1Arabic NLPSpeech Processing

  19. Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

    May 7, 2026Julius Richter, Yoshiki Masuyama, Christoph Boeddeker +3Stochastic InterpolantsSpeech Processing

  20. Spoken Language Identification with Pre-trained Models and Margin Loss

    May 3, 2026Zhihua Fang, Liang He, Weiwu JiangLanguage IdentificationSpeech Processing

  21. Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI

    May 2, 2026Yi-Cheng Lin, Yun-Shao Tsai, Kuan-Yu Chen +3Algorithmic FairnessSpeech Processing

  22. Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

    May 1, 2026Shakeel Sheikh, Patrick Marmaroli, MD Sahidullah +4Speech ProcessingHuman-in-the-Loop AI

  23. Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation

    May 1, 2026Anton Ratnarajah, Mehmet Ergezer, Arun Nair +1Synthetic Data AugmentationSpeech Processing

  24. Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints

    Apr 30, 2026Yurii Halychanskyi, Jianfeng Steven Guo, Volodymyr KindratenkoSpeech ProcessingVoice Conversion

  25. A Toolkit for Detecting Spurious Correlations in Speech Datasets

    Apr 29, 2026Lara Gauder, Pablo Riera, Andrea Slachevsky +3Speech ProcessingSpurious Correlation Robustness

  26. WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

    Apr 28, 2026Erfan Ramezani, Mohammad Mahdi Giahi, Mohammad Erfan Zarabadipour +2Streaming ASRSpeech Processing