eess.ASOct 7, 2026

Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification

Authors: Minu Kim, Jihwan Lee, David R. Mortensen, Shrikanth Narayanan

Organizations: Signal Analysis and Interpretation Lab (SAIL), University of Southern California, USA · Language Technologies Institute, Carnegie Mellon University, USA

Abstract

Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.

Figures & tables

Explore similar work

CardsList
  1. Contrastive Regularization for Accent-Robust ASR

    May 5, 2026Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham +1Contrastive LearningContrastive Loss

  2. Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?

    Oct 5, 2026Paula A. Perez-Toro, Tomas Arias-Vergara, Annette Schwarz +4