eess.ASOct 1, 2026

Code-Switching Spoken Language Identification as Multi-Label Set Prediction

Authors: Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima, Shinji Watanabe

Organizations: Waseda University, Tokyo, Japan · Johns Hopkins University, Baltimore, USA · Carnegie Mellon University, Pittsburgh, USA

Abstract

Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.

Figures & tables

Explore similar work

CardsList
  1. CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech

    Jun 9, 2026Brian Yan, Qingzheng Wang, Matthew Wiesner +9Code-SwitchingMultilingual Automatic Speech Recognition

  2. Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition

    Jun 5, 2026Tung X. Nguyen, Hieu Minh Truong, Giang Son Nguyen +3Code-SwitchingAutomatic Speech Recognition

  3. Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

    Jun 4, 2026Gio Paik, Hyunseo Shin, Soungmin LeeMultilingual Automatic Speech RecognitionAutomatic Speech Recognition