Speech Processing

Latest papers 295

All topics
CardsList
  1. Foreground Voice Activity Detection: Learning Speaker Selectivity from Supervision

    Sep 17, 2026Guangzhao Yang, Muhammad Huzaifah, Yu Pan +2Data AugmentationSpeech Processing

  2. A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech

    Sep 16, 2026Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi +4Cross-Lingual Representation AlignmentSpeech Processing

  3. GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

    Sep 16, 2026Zitao Liang, Chang GaoTTS SynthesisSpeech Processing

  4. OpenEnded: An Open-Response Speech Corpus for Speaking Proficiency Assessment with Human Annotations and ALM Supervision

    Sep 14, 2026Yu-Wen Chen, Eric Zhou, Evelyn Ding +3Speech Processing

  5. Interpreting hierarchical organisation of speaker embeddings

    Sep 14, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleySpeech ProcessingNeural Network Interpretability

  6. PhaseGAN: High-Fidelity Vocoder via Decoupled Amplitude and GAN-Driven Phase Reconstruction

    Sep 14, 2026Wenzheng Zhang, Xueliang Zhang, Shulin He +7Speech ProcessingAudio Generation

  7. Exploring Second-Order Pattern Recognition in Speaker Recognition

    Sep 12, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleyExplainable Artificial IntelligenceSpeech Processing

  8. ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

    Sep 12, 2026Luca Della Libera, Cem Subakan, Mirco RavanelliAudio Token CompressionNeural Audio Codecs

  9. Domain-Incremental Learning for Multi-Channel Replay Speech Detection

    Sep 11, 2026Michael NeriSpeech ProcessingDomain-Incremental Learning

  10. EConv-TasNet: Efficient Conv-TasNet for Effective Speech Separation

    Sep 11, 2026Pei-Chun Chang, Chuan-Yi LiuEfficient Neural Network InferenceTemporal Convolutional Networks

  11. Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

    Sep 10, 2026Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia +1Speech Generation EvaluationSpeech Processing

  12. Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

    Sep 9, 2026Soumyadeep RoyAudio Deepfake DetectionSpeech Processing

  13. Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

    Sep 9, 2026Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis +1TTS SynthesisSpeech Processing

  14. Source-Adaptive Data Curation for Bilingual NVV-Aware ASR

    Sep 9, 2026Yuang Cao, Qirui Zhan, Jingbin Hu +7Speech ProcessingAudio Understanding

  15. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

    Sep 9, 2026Hong Nguyen, Sean Foley, Christina Hagedorn +4Unsupervised Domain AdaptationJoint-Embedding Predictive Architecture

  16. SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

    Sep 9, 2026Jingyi Liao, Wenyu Zhang, Zhuohan Liu +6Audio-Language Model EvaluationASR Evaluation

  17. PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization

    Sep 3, 2026Fei Liu, Yang Ai, Xiao-Hang Jiang +1Neural Audio CodecsSpeech Processing

  18. Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

    Aug 13, 2026Serli Kopar, Sam Gijsen, Abner Hernandez +2Speech ProcessingParkinson's Disease

  19. myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

    Aug 11, 2026Ye Kyaw Thu, Ye Bhone Lin, Thura Aung +6Speech ProcessingAutomatic Speech Recognition

  20. Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

    Aug 10, 2026Abner Hernandez, Tomás Arias Vergara, Daiqi Liu +2Speech Processing

  21. Detection of Self-Introductions in Legislative Testimony

    Aug 8, 2026Sofija Dimitrijevic, Pallavi Das, Kasey Liu +1Speech Processing