Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.
Figures & tables
Subset
Matrix languages
Embedded languages
Read
ara deu fra hin por rus spa ces cmn ita jpn kor slk tel
eng
XTTS1
ara deu fra hin por rus spa ces cmn ita jpn kor hun nld pol tur
eng
XTTS2
ara cmn hin spa
ara fra deu ita por pol tur rus nld ces spa cmn jpn hun kor hin
MMS
ara deu fra hin por rus spa nld pol tur ben bul cat ceb cym ell fin guj heb hun ind isl jav kan kaz kir lav lug mal mar mya pan ron swe swh tam tel tgk tgl tha ukr urd uzb yor zlm
eng
TABLE I: Matrix and embedded language inventories of CS-FLEURS subsets used in this work. Language codes follow ISO 639-3.
ID
Model family
FLEURS
CS-FLEURS
CS-YODAS
Target form
H1
Hard-target multi-class classifier
√
–
–
Individual language class
H2
Hard-target multi-class classifier
√
√
–
Individual language class / atomic pair class
H3
Hard-target multi-class classifier
√
√
√
Individual language class / atomic pair class
S1
Soft-target classifier
√
√
–
Soft distribution over individual languages
S2
Soft-target classifier
√
√
√
Soft distribution over individual languages
M1
Sigmoid multi-label classifier
√
√
–
Binary multi-hot vector over individual languages
TABLE II: Summary of model configurations. Check marks indicate the datasets used for training.
ID
Model family
#Params
FL
CS-FL
CS-YO
FLEURS
CS-FLEURS
All
All
Seen
Unseen
H1-Top1 †
Hard-target multi-class classifier
9.33M
√
–
–
96.7
–
–
–
H1-Top2 †
Hard-target multi-class classifier
9.33M
√
–
–
–
8.3
10.7
6.7
S1-Top1 †
Soft-target classifier
9.32M
√
√
–
91.4
–
–
–
S1-Top2 †
Soft-target classifier
9.32M
√
√
–
–
52.0
95.2
22.2
S2-Top1 †
Soft-target classifier
9.32M
√
√
√
88.2
–
–
–
TABLE III: Exact set accuracy (%) on FLEURS and CS-FLEURS. FL, CS-FL, and CS-YO denote the training datasets FLEURS, CS-FLEURS, and CS-YODAS. † marks oracle Top- k ( k=1 for FLEURS and k=2 for CS-FLEURS). #Params excludes the shared 965M-parameter MMS-1B frontend. “–” means not applicable. Bold marks column maxima, including oracle results.
ID
Model family
#Params
FL
CS-FL
CS-YO
Read
XTTS1
XTTS2
MMS
Seen
Unseen
Seen
Unseen
Seen
Unseen
H1-Top2 †
Hard-target multi-class classifier
9.33M
√
–
–
14.6
1.2
1.7
0.5
33.7
16.3
S1-Top2 †
Soft-target classifier
9.32M
√
√
–
82.7
1.6
99.7
0.0
99.8
56.7
S2-Top2 †
Soft-target classifier
9.32M
√
√
√
5.7
1.0
77.2
0.0
84.5
37.0
M1-Top2 †
Sigmoid multi-label classifier
9.32M
√
√
–
88.1
1.2
99.6
0.0
99.9
87.3
M2-Top2 †
Sigmoid multi-label classifier
9.32M
√
√
√
88.6
19.0
99.6
0.0
99.9
82.0
TABLE IV: Subset-wise exact set accuracy (%) on CS-FLEURS. Seen and Unseen denote whether each language pair appears in CS training data. † marks oracle Top-2. Other notation follows Table III .
ID
Model family
#Params
FL
CS-FL
CS-YO
Read Seen
Read Unseen
XTTS1 Seen
XTTS2 Unseen
MMS Seen
MMS Unseen
H2
Hard-target multi-class classifier
9.33M
√
√
–
0.2
0.0
99.9
99.9
99.9
38.5
H3
Hard-target multi-class classifier
9.33M
√
√
√
0.8
0.0
99.8
99.7
99.9
14.6
G1
Multi-label set generator (ours)
23.55M
√
√
–
0.0
0.0
100.0
100.0
99.9
96.8
G2
Multi-label set generator (ours)
23.55M
√
√
√
0.2
0.0
100.0
99.9
99.7
43.1
TABLE V: Two-language prediction rate (%) on CS-FLEURS. For hard-target classifiers, atomic-pair predictions count as two-language predictions; for set generators, the rate is computed from unique generated language tokens. Notation follows Table III .
Fig. 1: Threshold sensitivity of the soft-target (left) and sigmoid (right) classifiers trained on FLEURS and CS-FLEURS (S1 and M1). Curves show exact set accuracy obtained by thresholding softmax probabilities or sigmoid scores; dotted lines show the corresponding oracle Top- k accuracies. No single threshold recovers Top- k accuracy across subsets.
English proportion
H2
S1-Top2 †
M1-Top2 †
G1
All
0.18
16.46
26.16
0.07
0–10%
0.00
57.83
70.87
0.00
10–20%
0.25
28.26
43.49
0.25
20–30%
0.00
20.65
32.59
0.00
30–40%
0.39
15.69
24.71
0.00
40–50%
0.00
17.65
25.88
0.00
TABLE VI: Exact set accuracy (%) on the Bangor Miami English–Spanish subset, grouped by English proportion. † marks oracle Top-2. Bold marks row maxima.
We present CS-YODAS, a Creative Commons-licensed dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching (CS), or the alternation between languages within an utterance or conversation, is common in multilingual settings but remains underrepresented in existing CS speech resources, which are typically small, domain-specific, or artificially constructed. Building on the YODAS corpus, we develop a scalable, human-in-the-loop pipeline for identifying and validating naturally occurring code-switching. The resulting dataset, which totals 313 hours and spans 7 matrix languages, provides diverse, real-world examples of spontaneous code-switched speech. We further analyze the distribution and characteristics of code-switching in the wild, examining language-pair frequencies and switching patterns, and report baseline results for spoken language identification. We hope that CS-YODAS will encourage broader and more comprehensive research on code-switched speech. Dataset link: https://huggingface.co/datasets/byan/cs-yodas.
Brian Yan, Qingzheng Wang, Matthew Wiesner +9
Carnegie Mellon University · Johns Hopkins University · University of Texas at Austin +4
Code-switching (CS), the alternation between multiple languages within a single utterance, remains challenging for Automatic Speech Recognition (ASR). To address this issue, we propose a Point-of-Interest (POI)-aware contrastive training framework that improves recognition at CS-critical regions. We first identify CS spans by adopting POI detection method from literature, then construct acoustically plausible near-miss hypotheses by perturbing POIs in ASR N-best outputs and expanding candidates with a large language model. Hard but plausible negatives are retained through filtering with acoustic, phonemic, and textual constraints. Finally, we fine-tune Whisper-small with LoRA using a POI-weighted cross-entropy anchor objective together with a multi-negative contrastive ranking loss. Experiments on CS-FLEURS (cmn-eng) and ViMedCSS (vie-eng) show consistent reductions of over 2% in both general and CS-aware error rates compared to standard LoRA fine-tuning.
Tung X. Nguyen, Hieu Minh Truong, Giang Son Nguyen +3
VinUniversity, Vietnam · University of Technology Sydney, Australia · Monash University, Australia
Automatic Speech Recognition (ASR) has become a key technology for human--AI interaction. However, code-switching ASR (CS-ASR) remains particularly challenging due to the severe scarcity of multilingual CS speech resources across diverse language pairs. Existing approaches primarily improve CS-ASR performance through synthetic CS speech generation or pair-specific fine-tuning on limited bilingual datasets. Nevertheless, these approaches face an inherent scalability limitation, as support for CS must be developed separately for language pairs whose number grows combinatorially with the number of supported languages. In this work, we investigate whether CS capabilities learned from a limited set of seen language pairs can generalize to unseen language pairs through model merging and domain generalization methods. Our experiments show that merged bilingual CS-ASR models modestly generalize to unseen language pairs, suggesting limited transfer of bilingual CS capabilities across language pairs.
Gio Paik, Hyunseo Shin, Soungmin Lee
Theta One Korea, Seoul, Republic of Korea · Kitsch Labs, Seongnam, Republic of Korea · Department of Artificial Intelligence, University of Seoul, Seoul, Republic of Korea +1