Speech Processing

Latest papers 295

All topics
CardsList
  1. FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching

    Mar 4, 2026Fabian Ritter-Gutierrez, Md Asif Jalal, Pablo Peso Parada +5Speech ProcessingVoice Conversion

  2. When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus

    Mar 2, 2026Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov +2Audio Deepfake DetectionSpeech Processing

  3. Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning

    Feb 24, 2026Bipasha Kashyap, Björn W. Schuller, Pubudu N. PathiranaRepresentation DisentanglementSpeech Processing

  4. Performance and Complexity Trade-off Optimization of Speech Models During Training

    Jan 20, 2026Esteban Gómez, Tom BäckströmModel CompressionNeural Network Optimization

  5. Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs

    Jan 18, 2026Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar ChowdhurySpoken Language UnderstandingSpeech Processing

  6. Aliasing-Free Neural Audio Synthesis

    Dec 23, 2025Yicheng Gu, Junan Zhang, Chaoren Wang +3Singing Voice SynthesisNeural Audio Codecs

  7. MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

    Dec 22, 2025Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun +2Speech ProcessingSelf-Supervised Speech Representation Learning

  8. Physics-Informed Neural Networks for Speech Production

    Nov 1, 2025Kazuya Yokota, Ryosuke Harakawa, Masaaki Baba +1Speech ProcessingPhysics-Informed ML

  9. UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

    Oct 23, 2025Haoyin Yan, Chengwei Liu, Shaofei Xue +4Target Speaker ExtractionDecoder-Only Language Models

  10. ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

    Oct 12, 2025Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh ShakeryTTS SynthesisSpeech Processing

  11. BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

    Sep 18, 2025Théo Charlot, Tarek Kunze, Maxime Poli +3Speech Foundation ModelsSpeech Processing

  12. An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment

    Aug 27, 2025Tien-Hong Lo, Szu-Yu Chen, Yao-Ting Sung +1Speech ProcessingOrdinal Regression

  13. Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

    Jul 17, 2025Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev +3Speech ProcessingSpeech Prosody

  14. Enforcing Speech Content Privacy in Environmental Sound Recordings using Segment-wise Waveform Reversal

    Jul 11, 2025Modan Tailleur, Mathieu Lagrange, Pierre Aumond +1Speech Processing

  15. Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker Traits and Speech Properties

    May 20, 2025Tiantian Feng, Jihwan Lee, Anfeng Xu +9Speech ProcessingSpeech Foundation Models

  16. BenSParX: A Robust Explainable Machine Learning Framework for Parkinson's Disease Detection from Bengali Conversational Speech

    May 18, 2025Riad Hossain, Muhammad Ashad Kabir, Arat Ibne Golam Mowla +2Speech ProcessingLow-Resource Language Processing

  17. Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption

    Mar 21, 2025Xiang Li, Pin-Yu Chen, Wenqi WeiAudio Deepfake DetectionSpeech Foundation Models

  18. Topology-enhanced machine learning for speech signal processing

    Nov 26, 2023Pingyao Feng, Qingrui Qu, Haiyu Zhang +4Interpretable MLSpeech Processing

  19. A mathematical model of the vowel space

    Oct 19, 2021Frédéric BerthommierSpeech Processing

  20. Personal VAD: Speaker-Conditioned Voice Activity Detection

    Aug 12, 2019Shaojin Ding, Quan Wang, Shuo-yiin Chang +2Speech Processing

  21. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    Date pendingTara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10Voice Agent EvaluationSpoken Dialogue Systems