Wav2Vec

Momentum

4 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 35

All topics
CardsList
  1. On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection

    Sep 30, 2026Lisan Al Amin, Lei Zhang, Vandana P. JanejaEnvironmental Sound Deepfake DetectionQuantum Kernels

  2. WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

    Sep 24, 2026Kwok-Ho Ng, Tingting Song, Bingwen Feng +1Audio Deepfake DetectionWav2Vec

  3. End-to-end Jordanian dialect speech-to-text self-supervised learning framework

    Sep 21, 2026Ali A. Safieh, Ibrahim Abu Alhaol, Rawan GhnematSelf-Supervised Speech ModelsArabic Natural Language Processing

  4. Do speech foundation models really learn words?

    Sep 9, 2026Robin Huo, Ewan DunbarSelf-Supervised Speech ModelsWav2Vec

  5. Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

    Aug 2, 2026Ilia Semenkov, Daria Kleeva, Ivan Dakhtin +2Wav2VecHuman Visual Cortex

  6. Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

    Jul 28, 2026Gyeongmin KimText-To-Speech SynthesisWav2Vec

  7. Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

    Jul 28, 2026Behrad TaghiBeyglou, Fatemeh Bagheri, Ervin SejdicDysarthric SpeechWav2Vec

  8. An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge

    Jul 14, 2026Shuming Fang, Shuifei ZengSpeaker DiarizationWav2Vec

  9. Why Do You Say It Like That? A Phoneme-level Framework for Explainable Speech Deepfake Detection

    Jul 9, 2026Anna Taylor, Michele Panariello, Massimiliano Todisco +3Audio Deepfake DetectionSpoofing

  10. InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

    Jul 7, 2026Samir Sadok, Xavier Alameda-PinedaSelf-Supervised Speech ModelsWav2Vec

  11. Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

    Jul 7, 2026Adrien Schneider, Kacper Zabkowski, Anderson Augusma +3Wav2VecAnonymization

  12. When and why do handcrafted cues help self-supervised anti-spoofing? A causal and faithfulness analysis

    Jul 5, 2026Yugwon WonSpoofingWav2Vec

  13. wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

    Jun 27, 2026James Tanner, Morgan Sonderegger, Jane Stuart-Smith +2Wav2VecGrapheme-To-Phoneme

  14. Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

    Jun 26, 2026Andrew C. Cullen, Neil G. Marchant, Jiani Xie +4AcousticWav2Vec

  15. What Does a Pathological Speech Assessment Model Know about Acoustic Features? A Case Study on Oral and Oropharyngeal Cancer Patients

    Jun 23, 2026Tuan Nguyen, Corinne Fredouille, Alain Ghio +2Wav2VecDysarthric Speech

  16. Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English

    Jun 22, 2026Hamid Mojarad, Kevin TangWav2VecGrapheme-To-Phoneme

  17. Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

    Jun 21, 2026Shuubham Ojha, Carol Espy-WilsonSpeech EnhancementWav2Vec

  18. How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

    Jun 20, 2026Abhijit Sinha, Hemant Kumar Kathania, Mohit Joshi +3Self-Supervised Speech ModelsWav2Vec

  19. Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

    Jun 20, 2026Jinghao Chen, Mostafa Shahin, Beena AhmedGrapheme-To-PhonemeWav2Vec

  20. Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

    Jun 18, 2026Paban Sapkota, Hemant Kumar Kathania, Sudarsana Reddy Kadiri +1Dysarthric SpeechData Augmentation

  21. Pretrained self-supervised speech models can recognize unseen consonants

    Jun 10, 2026Chihiro Taguchi, Éric Le Ferrand, Hirosi Nakagawa +4Self-Supervised Speech ModelsConsonants

  22. Towards Robust Arabic Speech Emotion Recognition with Deep Learning

    Jun 9, 2026Youcef Soufiane Gheffari, Samiya SilarbiArabic Natural Language ProcessingEmotion Recognition

  23. From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

    Jun 7, 2026Moshe Mandel, Shlomo E. ChazanVoice ConversionWav2Vec

  24. Building Community-Centred NLP Resources for Puno Quechua

    May 27, 2026Elwin Huaman, Adrian Gamarra Lafuente, Johanna Cordova +1Low-Resource LanguagesAutomatic Speech Recognition Evaluation

  25. Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

    May 9, 2026Aishwarya Fursule, Shruti Kshirsagar, Anderson R. AvilaEnvironmental Sound Deepfake DetectionWav2Vec

  26. Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

    May 4, 2026Sandra Arcos-Holzinger, Sarah M. Erfani, James Bailey +1Self-Supervised Speech ModelsWav2Vec

  27. Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection

    Apr 28, 2026Jaskirat Sudan, Hashim Ali, Surya Subramani +1Audio Deepfake DetectionWav2Vec

  28. Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0

    Apr 23, 2026Natalie Engert, Dominik Wagner, Korbinian Riedhammer +1Wav2VecDysarthric Speech

  29. DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

    Apr 23, 2026Nikhil RaghavSpeaker DiarizationSpeaker

  30. Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages

    Apr 20, 2026V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard JägerGrapheme-To-PhonemeWav2Vec

  31. Echoes: A semantically-aligned music deepfake detection dataset

    Mar 24, 2026Octavian Pascu, Dan Oneata, Horia Cucu +1Audio Deepfake DetectionDeepfake Detection

  32. The NTNU System at the S&I Challenge 2025 SLA Open Track

    Jun 5, 2025Hong-Yun Lin, Tien-Hong Lo, Yu-Hsuan Fang +4Wav2VecTranscript