Self-Supervised Speech Representation Learning

Latest papers 70

All topics
CardsList
  1. Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English

    Jun 22, 2026Hamid Mojarad, Kevin TangSelf-Supervised Speech Representation Learning

  2. How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

    Jun 20, 2026Abhijit Sinha, Hemant Kumar Kathania, Mohit Joshi +3Speech ProcessingSelf-Supervised Speech Representation Learning

  3. Online Predictive Coding for Dual-Mode Self-Supervised Speech Model

    Jun 19, 2026Keita Goto, Takashi Maekaku, Jin Sakuma +3Predictive CodingStreaming ASR

  4. A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

    Jun 18, 2026Nabil Mosharraf Hossain, Riasat Islam, Unaizah ObaidellahFine-TuningAutomatic Speech Recognition

  5. S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

    Jun 17, 2026Georgios Ioannides, Adrian Kieback, Judah Goldfeder +5ClusteringJoint-Embedding Predictive Architecture

  6. Perceptual compensation for tonal context in self-supervised speech models

    Jun 16, 2026James Kirby, Ioana Krehan, Michele GubianSupervised Fine-TuningSpeech Prosody

  7. SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

    Jun 15, 2026Sejal Bhalla, Larry Kieu, Aina Merchant +2Speech ProcessingSelf-Supervised Speech Representation Learning

  8. ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

    Jun 15, 2026Zeqian Hu, Fuliang Weng, Shu Shang +1Zero-Shot LearningJoint-Embedding Predictive Architecture

  9. From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

    Jun 11, 2026Pedro Correa, Olivier Perrotin, Samir Sadok +2Audio Representation LearningNeural Audio Codecs

  10. From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation

    Jun 11, 2026Fengrui Liu, Ruiyang Huang, Qijian Zheng +2Synthetic Data PretrainingAudio Representation Learning

  11. Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

    Jun 10, 2026Chiara Semenzin, Faadil Mustun, Roberto Dessi +5Audio Representation LearningAudio Understanding

  12. Pretrained self-supervised speech models can recognize unseen consonants

    Jun 10, 2026Chihiro Taguchi, Éric Le Ferrand, Hirosi Nakagawa +4Automatic Speech RecognitionSelf-Supervised Speech Representation Learning

  13. Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

    Jun 9, 2026Zilai Wang, Natarajan Balaji Shankar, Mohan Shi +2Domain Adaptation for ASRDomain Adaptation

  14. Speaker Group Encoding in Self-supervised Speech Recognition Models

    Jun 9, 2026Felix Herron, Solange Rossato Alexandre Allauzen, Benoit Favre +1Audio Representation LearningDemographic Bias in ASR

  15. ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

    Jun 9, 2026Khanh Le, Kiet Anh Hoang, Bao Nguyen +5Audio Representation LearningSpeech Processing

  16. Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

    Jun 8, 2026Diane Myung-kyung Woodbridge, Jee Hyun SuhSpeech ProcessingSelf-Supervised Speech Representation Learning

  17. A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

    Jun 7, 2026Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier +1Audio Deepfake DetectionDomain Generalization

  18. Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

    Jun 5, 2026Kentaro Onda, Satoru Fukayama, Daisuke Saito +1Automatic Speech RecognitionAudio Tokenization

  19. Context-aware child-directed speech detection from long-form recordings

    May 31, 2026Théo Charlot, Tarek Kunze, Kaveri K. Sheth +2Speech ProcessingAudio Understanding

  20. Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

    May 26, 2026Serli Kopar, Roshan Prakash Rane, Christian Mychajliw +6Speech ProcessingAudio Understanding

  21. MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

    May 22, 2026Qingcao Li, Yipeng Lin, Weichen Lian +3Audio Deepfake DetectionPrompt Tuning

  22. Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

    May 13, 2026Wuao Liu, Mustafa Chasmai, Subhransu Maji +1Masked AutoencodersBioacoustics

  23. PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization

    May 7, 2026Adhiraj Banerjee, Vipul AroraAudio TokenizationSelf-Supervised Speech Representation Learning

  24. WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

    May 7, 2026Guanrou Yang, Tian Tan, Qian Chen +12Audio Representation LearningTTS Synthesis

  25. A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language

    May 5, 2026Thibault Bañeras-Roux, Mickael Rouvier, Jane Wottawa +1ASR EvaluationAutomatic Speech Recognition

  26. Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

    May 4, 2026Sandra Arcos-Holzinger, Sarah M. Erfani, James Bailey +1Intrinsic DimensionalityUnsupervised Anomaly Detection

  27. Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

    Apr 24, 2026Felix Herron, Solange Rossato, Alexandre Allauzen +1Demographic Bias in ASRAutomatic Speech Recognition

  28. Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus

    Apr 24, 2026Szu-Jui Chen, John H. L. HansenAutomatic Speech RecognitionSelf-Supervised Speech Representation Learning