Speech Processing

Latest papers 295

All topics
CardsList
  1. ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

    Apr 28, 2026Kexue Wang, Yinfeng Yu, Liejun WangSpeech ProcessingEmotion Recognition in Conversations

  2. Keyword spotting using convolutional neural network for speech recognition in Hindi

    Apr 26, 2026Saru Bharti, Pushparaj Mani PathakKeyword SpottingSpeech Processing

  3. Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

    Apr 25, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleyClusteringHierarchical Clustering

  4. Au-M-ol: A Unified Model for Medical Audio and Language Understanding

    Apr 25, 2026Meizhu Liu, Nistha Mitra, Paul Li +3Speech ProcessingAudio Understanding

  5. Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection

    Apr 25, 2026Khalid Zaman, Masashi UnokiAudio Deepfake DetectionSpeech Processing

  6. Prompting Whisper for Joint Speech Transcription and Diarization

    Apr 24, 2026Mariia Zamyrova, Henk van den HeuvelSpeech ProcessingAutomatic Speech Recognition

  7. Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

    Apr 24, 2026Haopeng Geng, Longfei Yang, Xi Chen +3Speech Processing

  8. Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0

    Apr 23, 2026Natalie Engert, Dominik Wagner, Korbinian Riedhammer +1Speech ProcessingSelf-Supervised Speech Representation Learning

  9. Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines

    Apr 22, 2026Hawau Olamide Toyin, Mutiah Apampa, Toluwani Aremu +6Speech ProcessingAutomatic Speech Recognition

  10. Enhancing Speaker Verification with Whispered Speech via Post-Processing

    Apr 22, 2026Magdalena Gołębiowska, Piotr SygaAudio Representation LearningSpeaker Verification

  11. StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

    Apr 21, 2026Shuhai Peng, Hui Lu, Jinjiang Liu +8Target Speaker ExtractionSpeech Processing

  12. Tadabur: A Large-Scale Quran Audio Dataset

    Apr 21, 2026Faisal AlherranSpeech Processing

  13. Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

    Apr 20, 2026Chenqian Le, Ruisi Li, Beatrice Fumagalli +6Speech Processing

  14. Where Do Self-Supervised Speech Models Become Unfair?

    Apr 20, 2026Felix Herron, Maja Hjuler, Solange Rossato +2Speech ProcessingDemographic Bias in ASR

  15. Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model

    Apr 16, 2026Adelekun Oluwademilade, Ademola Adedamola, Abiola Abdulhakeem +7Speech ProcessingSpeech Emotion Recognition

  16. ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning

    Apr 16, 2026Junyi Wang, Chi Zhang, Jing Qian +4Neural Audio CodecsSpeech Processing

  17. The Acoustic Camouflage Phenomenon: Re-evaluating Speech Features for Financial Risk Prediction

    Apr 16, 2026Dhruvin Dungrani, Disha DungraniSpeech ProcessingMultimodal Learning

  18. DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio

    Apr 10, 2026Wataru Nakata, Yuki Saito, Kazuki Yamauchi +2Audio Diffusion ModelsSpeech Processing

  19. KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness

    Mar 30, 2026Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo +4Audio-Language Model EvaluationSpeech Processing

  20. PHONOS: PHOnetic Neutralization for Online Streaming Applications

    Mar 27, 2026Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah +1Speech ProcessingSpeaker Anonymization

  21. SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation

    Mar 23, 2026Lucas H. Ueda, João G. T. Lima, Pedro R. Corrêa +3Disentangled Representation LearningTTS Synthesis

  22. Voice Privacy from an Attribute-based Perspective

    Mar 19, 2026Mehtab Ur Rahman, Martha Larson, Cristian Tejedor-GarciaPrivacy AuditingSpeech Processing

  23. DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

    Mar 19, 2026Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo +3Speech ProcessingSelf-Supervised Speech Representation Learning

  24. Controllable Accent Normalization via Discrete Diffusion

    Mar 15, 2026Qibing Bai, Yuhan Du, Tom Ko +3Audio Diffusion ModelsSpeech Processing

  25. What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection

    Mar 14, 2026Shree Harsha Bokkahalli Satish, Harm Lameris, Joakim Gustafson +1Audio Deepfake DetectionSpeech Processing

  26. Mask2Flow-TSE: Two-Stage Target Speaker Extraction with Masking and Flow Matching

    Mar 13, 2026Junwon Moon, Seungbeom Kim, Hansol Park +4Target Speaker ExtractionFlow Matching

  27. Speech Codec Probing from Semantic and Phonetic Perspectives

    Mar 11, 2026Xuan Shi, Chang Zeng, Tiantian Feng +3Audio Representation LearningNeural Audio Codecs

  28. Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks

    Mar 9, 2026Pol Buitrago, Oriol Pareras, Federico Costa +1Speaker VerificationSpeech Processing

  29. PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment

    Mar 9, 2026Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1Benchmark DesignSpeech Quality Assessment