Self-Supervised Speech Representation Learning

Latest papers 70

All topics
CardsList
  1. Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification

    Oct 7, 2026Minu Kim, Jihwan Lee, David R. Mortensen +1Language IdentificationAlgorithmic Bias

  2. Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?

    Oct 5, 2026Paula A. Perez-Toro, Tomas Arias-Vergara, Annette Schwarz +4Audio Representation LearningSelf-Supervised Speech Representation Learning

  3. Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

    Oct 1, 2026Serli Kopar, Alkis Koudounas, Roshan P. Rane +3Audio Representation LearningSpeech-Based Dementia Detection

  4. GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

    Sep 29, 2026Gaspard Botté, Séverin Baroudi, Samir Sadok +5Joint-Embedding Predictive ArchitecturePredictive Representation Learning

  5. GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

    Sep 28, 2026Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang +1Audio Deepfake DetectionDomain Generalization

  6. OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations

    Sep 28, 2026Faadil Mustun, Chiara Semenzin, Roberto Dessi +7Sound Event DetectionBioacoustics

  7. SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations

    Sep 28, 2026Xiaoyu Yang, Arthur Hinsvark, Antonios Alexos +3Representation LearningSelf-Supervised Speech Representation Learning

  8. What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

    Sep 27, 2026Jiajun Xu, Menglu Li, Xiao-Ping ZhangAudio Representation LearningAudio Deepfake Detection

  9. Do Audio Language Models Hear and Read Distinctive Features Alike?

    Sep 24, 2026Yuanhao Chen, Peter ChinAudio-Language ModelsSelf-Supervised Speech Representation Learning

  10. ToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification

    Sep 21, 2026Qisheng Liao, Youngah DoContrastive LearningFew-Shot Audio Classification

  11. End-to-end Jordanian dialect speech-to-text self-supervised learning framework

    Sep 21, 2026Ali A. Safieh, Ibrahim Abu Alhaol, Rawan GhnematAutomatic Speech RecognitionSelf-Supervised Speech Representation Learning

  12. Do speech foundation models really learn words?

    Sep 9, 2026Robin Huo, Ewan DunbarRepresentation LearningSpeech Foundation Models

  13. Is Semantics Enough for Speech Mean Opinion Score Prediction?

    Sep 3, 2026Tianyu Lan, Yufei Shi, Yang Ai +3Human Preference EvaluationNeural Audio Codecs

  14. Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

    Aug 19, 2026Wenxuan He, Yunpeng Li, Zewei Li +4Soft-Label LearningUnsupervised Clustering

  15. Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

    Aug 13, 2026Serli Kopar, Sam Gijsen, Abner Hernandez +2Speech ProcessingParkinson's Disease

  16. Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

    Jul 28, 2026Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3Audio Representation LearningNeural Audio Codecs

  17. Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

    Jul 28, 2026Behrad TaghiBeyglou, Fatemeh Bagheri, Ervin SejdicGraph Representation LearningClinical Outcome Prediction

  18. Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

    Jul 21, 2026Laurin Wagner, Bernhard Thallinger, Miroslav Stankovic +1Audio TokenizationSelf-Supervised Speech Representation Learning

  19. Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?

    Jul 15, 2026Wangjin Zhou, Yizhou Zhang, Yichi Wang +1Supervised Fine-TuningFine-Tuning

  20. GigaAM Multilingual: Foundation Model for Underrepresented Languages

    Jul 11, 2026Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov +5Speech Foundation ModelsSpeech Processing

  21. When and why do handcrafted cues help self-supervised anti-spoofing? A causal and faithfulness analysis

    Jul 5, 2026Yugwon WonAudio Deepfake DetectionASVspoof

  22. Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

    Jul 5, 2026Ryota Komatsu, Kota Kawakita, Takuma Okamoto +1Audio TokenizationSelf-Supervised Speech Representation Learning

  23. How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

    Jun 30, 2026Ailín Pollio San Pedro, Tomi Kinnunen, Alexandre Nikolaev +1Speech ProcessingSelf-Supervised Speech Representation Learning

  24. BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

    Jun 29, 2026Ludovic K. Tuncay, Etienne Labbé, Thomas PellegriniAudio Representation LearningSelf-Supervised Pre-Training

  25. Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

    Jun 24, 2026Jingjing Xu, Zijian Yang, Mohammad Zeineldeen +3Automatic Speech RecognitionVector Quantization