Speech Foundation Models

Latest papers 42

All topics
CardsList
  1. Backdooring Acoustic Foundation Models for Physically Realizable Triggers

    Oct 7, 2026Zebin Yun, Eyal Ronen, Mahmood SharifBackdoor AttacksAdversarial Attacks

  2. Neural Representations, Natural Connections: What Transfers From Human Speech Foundation Models to Animal Vocalizations?

    Oct 5, 2026Tomás Arias-Vergara, Christopher Hauer, Héloïse Brotier +3Audio Representation LearningSpeech Foundation Models

  3. Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization

    Oct 4, 2026Aref Mousavi, Shahab Sherafat, Kiarash Kiani Feriz +4Adversarial Representation LearningSpeech Foundation Models

  4. Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition

    Sep 24, 2026Asmee Mishra, Mengjie Qian, Brechtje Post +1Speech Foundation ModelsLow-Resource Speech Recognition

  5. Temporal Taxation Compounds Under Post-Training Compression of Whisper Models

    Sep 23, 2026Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers +1ASR EvaluationSpeech Foundation Models

  6. Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

    Sep 23, 2026Rasmus Aagaard, Nicki Skafte DetlefsenSpeech Foundation ModelsAutomatic Speech Recognition

  7. Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

    Sep 21, 2026Lujia Bao, Qian Chen, Luyao Cheng +15Tool-Augmented Language Model AgentsSpeech Foundation Models

  8. AURA: Uncertainty-Routed Activation Editing for Acoustic Grounding in Speech Foundation Models

    Sep 21, 2026Natarajan Balaji Shankar, Zilai Wang, Zihan Wang +3Speech Foundation Models

  9. Listening for Airway Stenosis: A Foundation Model-Based Method for Rapid and Accessible Detection

    Sep 14, 2026Jean Groeninger, Zihao Zhao, Juliana de Castilhos +2Speech Foundation ModelsAudio Classification

  10. CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models

    Sep 14, 2026Alef Iury Siqueira Ferreira, Pedro Lustosa Rege Botelho, Fernanda Silva +5Speech Foundation ModelsSpeech Quality Assessment

  11. Sparse Weight and Edge Circuit Discovery in Transformer-based Acoustic Models

    Sep 11, 2026Jiankun Wei, Ewan Dunbar, Gerald PennTransformer InterpretabilityCircuit Discovery

  12. Do speech foundation models really learn words?

    Sep 9, 2026Robin Huo, Ewan DunbarRepresentation LearningSpeech Foundation Models

  13. BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

    Sep 9, 2026Shivam Singh, Aditya Yadavalli, Catherine Arnett +1Speech Foundation ModelsAutomatic Speech Recognition

  14. Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection

    Sep 7, 2026Maryam Abbasihafshejani, Murtuza JadliwalaSpeech Foundation ModelsAutomatic Speech Recognition

  15. Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

    Aug 12, 2026Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan +4Speech Foundation ModelsDomain Generalization

  16. MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

    Aug 5, 2026Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen +18Speech Foundation ModelsSpeech Processing

  17. dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

    Aug 2, 2026Hankun Wang, Bohan Li, Shi Lian +7Speech Foundation ModelsSpeech Editing

  18. Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition

    Jul 31, 2026Susmita Bhattacharjee, Jagabandhu Mishra, H. S. Shekhawat +2Speech Foundation ModelsMeta-Learning

  19. Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage

    Jul 27, 2026Vivek Senthil, Ernest FokouéSpeech Foundation ModelsAudio Understanding

  20. Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

    Jul 26, 2026Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai +1Zero-Shot LearningSpeech Foundation Models

  21. Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?

    Jul 15, 2026Wangjin Zhou, Yizhou Zhang, Yichi Wang +1Supervised Fine-TuningFine-Tuning

  22. GigaAM Multilingual: Foundation Model for Underrepresented Languages

    Jul 11, 2026Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov +5Speech Foundation ModelsSpeech Processing

  23. wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

    Jun 27, 2026James Tanner, Morgan Sonderegger, Jane Stuart-Smith +2Speech Foundation ModelsSpeech Processing

  24. Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

    Jun 10, 2026Haoning Xu, Zhaoqing Li, Huimeng Wang +4Speech Foundation ModelsClustering