Speech Processing

Latest papers 293

All topics
CardsList
  1. A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization

    Oct 6, 2026Tien-Hong Lo, Fong-Chun Tsai, Ting-An Hung +3Speech ProcessingSpeech Prosody

  2. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

    Oct 4, 2026Abdul Rehman, Jian-Jun Zhang, Xiaosong YangTTS SynthesisSpeech Processing

  3. SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic

    Oct 1, 2026Ben Sapirstein, Roy Mattar, Guy Mor-Lan +3ASR EvaluationArabic NLP

  4. Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

    Oct 1, 2026Hwayeon Kim, Youngwon Choi, Hyeonyu KimModel CompressionSpeech Processing

  5. Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

    Oct 1, 2026Abner Hernandez, Tomás Arias Vergara, Andreas Maier +1Speech ProcessingChild Language Acquisition

  6. NEUROTOKEN: Joint Source and Directional AAD with Envelope Decoding via Conditional Flow Matching

    Sep 30, 2026Ali Alavi, Donald S. WilliamsonNeural DecodingEEG Decoding

  7. VOSSA: Voiceprint Optimization for Streaming Speech Architectures

    Sep 30, 2026Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah +1Audio Representation LearningSpeech Processing

  8. Towards Model as a Library: Offline, Community-Sourced AI for Low-Resource African Languages

    Sep 29, 2026Fendji K. E. Jean LouisSpeech ProcessingLow-Resource Language Processing

  9. Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

    Sep 29, 2026Joseph T Colonel, Daniel Katzman, Kelsey Kirker +6Audio-Language Model EvaluationSpeaker Verification

  10. Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device

    Sep 28, 2026Paweł Warlewski, Artur Czeczko, Artur Szumaczuk +2Keyword SpottingSpeech Processing

  11. Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents

    Sep 28, 2026Suhwan Choi, Myeongho Jeon, Myungjoo KangSpeech Processing

  12. Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

    Sep 24, 2026Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves LaprieSpeech Processing

  13. Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

    Sep 24, 2026Michael Neri, Archontis Politis, Tuomas VirtanenSim-to-Real TransferSpeech Processing

  14. BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

    Sep 24, 2026Nikhil Navas, Sergio Chevtchenko, Talisson Damiao +1Speech ProcessingLow-Resource Speech Recognition

  15. Towards clinical adoption of voice and speech as measures of health: the need for harmonization

    Sep 24, 2026Nicholas Cummins, Vikram Ramanarayanan, Daniel Low +9Speech ProcessingClinical Speech Processing

  16. ASR ensembling for phoneme intelligibility evaluation of speech anonymizers

    Sep 23, 2026Victor Ménestrel, Sebastian Möller, Slim Ouni +1ASR EvaluationSpeech Processing

  17. Quieter Than the Room: Representation Drift and Task Robustness in Speech Encoders

    Sep 23, 2026Vsevolod Kovalev, Pranay ManochaSpeech Processing

  18. NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

    Sep 22, 2026Peter Sullivan, Bashar Talafha, Ahmed Ashraf +11ASR EvaluationSpoken Language Understanding

  19. Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

    Sep 22, 2026Robert Sutherland, Stefan Goetze, Jon BarkerTarget Speaker ExtractionSpeech Processing

  20. Narrowband Voice Communication Using Streaming Neural Compression

    Sep 21, 2026Dahong Luo, Anannya Trehan, Aritrik Ghosh +1Efficient Neural Network InferenceNeural Audio Codecs

  21. Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

    Sep 21, 2026João Lima, Lucas Ueda, Paula CostaSpeech ProcessingSpeech Prosody

  22. Understanding Hyperspherical Geometry of ECAPA-TDNN Embedding and Its Impact on Zero-Shot Voice Conversion

    Sep 21, 2026Mathilde Abrassart, Nicolas Obin, Axel RoebelNeural Representation GeometrySpeaker Verification

  23. Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis

    Sep 21, 2026Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4TTS SynthesisAudio Reasoning

  24. P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution

    Sep 21, 2026Ningyuan Yang, Yize Li, Pu Zhao +5Flow MatchingSpeech Processing

  25. ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

    Sep 21, 2026Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1Speech Quality AssessmentSpeech Processing

  26. Foreground Voice Activity Detection: Learning Speaker Selectivity from Supervision

    Sep 17, 2026Guangzhao Yang, Muhammad Huzaifah, Yu Pan +2Data AugmentationSpeech Processing

  27. A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech

    Sep 16, 2026Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi +4Cross-Lingual Representation AlignmentSpeech Processing

  28. GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

    Sep 16, 2026Zitao Liang, Chang GaoTTS SynthesisSpeech Processing

  29. OpenEnded: An Open-Response Speech Corpus for Speaking Proficiency Assessment with Human Annotations and ALM Supervision

    Sep 14, 2026Yu-Wen Chen, Eric Zhou, Evelyn Ding +3Speech Processing

  30. Interpreting hierarchical organisation of speaker embeddings

    Sep 14, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleySpeech ProcessingNeural Network Interpretability

  31. PhaseGAN: High-Fidelity Vocoder via Decoupled Amplitude and GAN-Driven Phase Reconstruction

    Sep 14, 2026Wenzheng Zhang, Xueliang Zhang, Shulin He +7Speech ProcessingAudio Generation

  32. Exploring Second-Order Pattern Recognition in Speaker Recognition

    Sep 12, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleyExplainable Artificial IntelligenceSpeech Processing

  33. ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

    Sep 12, 2026Luca Della Libera, Cem Subakan, Mirco RavanelliAudio Token CompressionNeural Audio Codecs

  34. Domain-Incremental Learning for Multi-Channel Replay Speech Detection

    Sep 11, 2026Michael NeriSpeech ProcessingDomain-Incremental Learning

  35. EConv-TasNet: Efficient Conv-TasNet for Effective Speech Separation

    Sep 11, 2026Pei-Chun Chang, Chuan-Yi LiuEfficient Neural Network InferenceTemporal Convolutional Networks

  36. Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

    Sep 10, 2026Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia +1Speech Generation EvaluationSpeech Processing

  37. Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

    Sep 9, 2026Soumyadeep RoyAudio Deepfake DetectionSpeech Processing

  38. Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

    Sep 9, 2026Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis +1TTS SynthesisSpeech Processing

  39. Source-Adaptive Data Curation for Bilingual NVV-Aware ASR

    Sep 9, 2026Yuang Cao, Qirui Zhan, Jingbin Hu +7Speech ProcessingAudio Understanding

  40. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

    Sep 9, 2026Hong Nguyen, Sean Foley, Christina Hagedorn +4Unsupervised Domain AdaptationJoint-Embedding Predictive Architecture

  41. SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

    Sep 9, 2026Jingyi Liao, Wenyu Zhang, Zhuohan Liu +6Audio-Language Model EvaluationASR Evaluation

  42. PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization

    Sep 3, 2026Fei Liu, Yang Ai, Xiao-Hang Jiang +1Neural Audio CodecsSpeech Processing

  43. Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

    Aug 13, 2026Serli Kopar, Sam Gijsen, Abner Hernandez +2Speech ProcessingParkinson's Disease

  44. myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

    Aug 11, 2026Ye Kyaw Thu, Ye Bhone Lin, Thura Aung +6Speech ProcessingAutomatic Speech Recognition

  45. Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

    Aug 10, 2026Abner Hernandez, Tomás Arias Vergara, Daiqi Liu +2Speech Processing

  46. Detection of Self-Introductions in Legislative Testimony

    Aug 8, 2026Sofija Dimitrijevic, Pallavi Das, Kasey Liu +1Speech Processing

  47. Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

    Aug 6, 2026Zezhong Jin, Xiaoyu Wang, Zhe Li +4Audio Representation LearningResidual Learning