Speech Processing

Latest papers 295

All topics
CardsList
  1. A multi-scenario EEG dataset for auditory attention decoding in naturalistic multi-talker environments

    Oct 7, 2026Shu Peng, Rui Liu, Yufei Zhang +7EEG DecodingSpeech Processing

  2. A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization

    Oct 6, 2026Tien-Hong Lo, Fong-Chun Tsai, Ting-An Hung +3Speech ProcessingSpeech Prosody

  3. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

    Oct 4, 2026Abdul Rehman, Jian-Jun Zhang, Xiaosong YangTTS SynthesisSpeech Processing

  4. SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic

    Oct 1, 2026Ben Sapirstein, Roy Mattar, Guy Mor-Lan +3ASR EvaluationArabic NLP

  5. Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

    Oct 1, 2026Hwayeon Kim, Youngwon Choi, Hyeonyu KimModel CompressionSpeech Processing

  6. Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

    Oct 1, 2026Abner Hernandez, Tomás Arias Vergara, Andreas Maier +1Speech ProcessingChild Language Acquisition

  7. NEUROTOKEN: Joint Source and Directional AAD with Envelope Decoding via Conditional Flow Matching

    Sep 30, 2026Ali Alavi, Donald S. WilliamsonNeural DecodingEEG Decoding

  8. VOSSA: Voiceprint Optimization for Streaming Speech Architectures

    Sep 30, 2026Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah +1Audio Representation LearningSpeech Processing

  9. Towards Model as a Library: Offline, Community-Sourced AI for Low-Resource African Languages

    Sep 29, 2026Fendji K. E. Jean LouisSpeech ProcessingLow-Resource Language Processing

  10. Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

    Sep 29, 2026Joseph T Colonel, Daniel Katzman, Kelsey Kirker +6Audio-Language Model EvaluationSpeaker Verification

  11. Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device

    Sep 28, 2026Paweł Warlewski, Artur Czeczko, Artur Szumaczuk +2Keyword SpottingSpeech Processing

  12. Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents

    Sep 28, 2026Suhwan Choi, Myeongho Jeon, Myungjoo KangSpeech Processing

  13. Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

    Sep 24, 2026Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves LaprieSpeech Processing

  14. Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

    Sep 24, 2026Michael Neri, Archontis Politis, Tuomas VirtanenSim-to-Real TransferSpeech Processing

  15. BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

    Sep 24, 2026Nikhil Navas, Sergio Chevtchenko, Talisson Damiao +1Speech ProcessingLow-Resource Speech Recognition

  16. Towards clinical adoption of voice and speech as measures of health: the need for harmonization

    Sep 24, 2026Nicholas Cummins, Vikram Ramanarayanan, Daniel Low +9Speech ProcessingClinical Speech Processing

  17. ASR ensembling for phoneme intelligibility evaluation of speech anonymizers

    Sep 23, 2026Victor Ménestrel, Sebastian Möller, Slim Ouni +1ASR EvaluationSpeech Processing

  18. Quieter Than the Room: Representation Drift and Task Robustness in Speech Encoders

    Sep 23, 2026Vsevolod Kovalev, Pranay ManochaSpeech Processing

  19. NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

    Sep 22, 2026Peter Sullivan, Bashar Talafha, Ahmed Ashraf +11ASR EvaluationSpoken Language Understanding

  20. Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

    Sep 22, 2026Robert Sutherland, Stefan Goetze, Jon BarkerTarget Speaker ExtractionSpeech Processing

  21. Narrowband Voice Communication Using Streaming Neural Compression

    Sep 21, 2026Dahong Luo, Anannya Trehan, Aritrik Ghosh +1Efficient Neural Network InferenceNeural Audio Codecs

  22. Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

    Sep 21, 2026João Lima, Lucas Ueda, Paula CostaSpeech ProcessingSpeech Prosody

  23. Understanding Hyperspherical Geometry of ECAPA-TDNN Embedding and Its Impact on Zero-Shot Voice Conversion

    Sep 21, 2026Mathilde Abrassart, Nicolas Obin, Axel RoebelNeural Representation GeometrySpeaker Verification

  24. Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis

    Sep 21, 2026Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4TTS SynthesisAudio Reasoning

  25. P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution

    Sep 21, 2026Ningyuan Yang, Yize Li, Pu Zhao +5Flow MatchingSpeech Processing

  26. ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

    Sep 21, 2026Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1Speech Quality AssessmentSpeech Processing