Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.
Figures & tables
Subset
Matrix languages
Embedded languages
Read
ara deu fra hin por rus spa ces cmn ita jpn kor slk tel
eng
XTTS1
ara deu fra hin por rus spa ces cmn ita jpn kor hun nld pol tur
eng
XTTS2
ara cmn hin spa
ara fra deu ita por pol tur rus nld ces spa cmn jpn hun kor hin
MMS
ara deu fra hin por rus spa nld pol tur ben bul cat ceb cym ell fin guj heb hun ind isl jav kan kaz kir lav lug mal mar mya pan ron swe swh tam tel tgk tgl tha ukr urd uzb yor zlm
eng
TABLE I: Matrix and embedded language inventories of CS-FLEURS subsets used in this work. Language codes follow ISO 639-3.
ID
Model family
FLEURS
CS-FLEURS
CS-YODAS
Target form
H1
Hard-target multi-class classifier
√
–
–
Individual language class
H2
Hard-target multi-class classifier
√
√
–
Individual language class / atomic pair class
H3
Hard-target multi-class classifier
√
√
√
Individual language class / atomic pair class
S1
Soft-target classifier
√
√
–
Soft distribution over individual languages
S2
Soft-target classifier
√
√
√
Soft distribution over individual languages
M1
Sigmoid multi-label classifier
√
√
–
Binary multi-hot vector over individual languages
TABLE II: Summary of model configurations. Check marks indicate the datasets used for training.
ID
Model family
#Params
FL
CS-FL
CS-YO
FLEURS
CS-FLEURS
All
All
Seen
Unseen
H1-Top1 †
Hard-target multi-class classifier
9.33M
√
–
–
96.7
–
–
–
H1-Top2 †
Hard-target multi-class classifier
9.33M
√
–
–
–
8.3
10.7
6.7
S1-Top1 †
Soft-target classifier
9.32M
√
√
–
91.4
–
–
–
S1-Top2 †
Soft-target classifier
9.32M
√
√
–
–
52.0
95.2
22.2
S2-Top1 †
Soft-target classifier
9.32M
√
√
√
88.2
–
–
–
TABLE III: Exact set accuracy (%) on FLEURS and CS-FLEURS. FL, CS-FL, and CS-YO denote the training datasets FLEURS, CS-FLEURS, and CS-YODAS. † marks oracle Top- k ( k=1 for FLEURS and k=2 for CS-FLEURS). #Params excludes the shared 965M-parameter MMS-1B frontend. “–” means not applicable. Bold marks column maxima, including oracle results.
ID
Model family
#Params
FL
CS-FL
CS-YO
Read
XTTS1
XTTS2
MMS
Seen
Unseen
Seen
Unseen
Seen
Unseen
H1-Top2 †
Hard-target multi-class classifier
9.33M
√
–
–
14.6
1.2
1.7
0.5
33.7
16.3
S1-Top2 †
Soft-target classifier
9.32M
√
√
–
82.7
1.6
99.7
0.0
99.8
56.7
S2-Top2 †
Soft-target classifier
9.32M
√
√
√
5.7
1.0
77.2
0.0
84.5
37.0
M1-Top2 †
Sigmoid multi-label classifier
9.32M
√
√
–
88.1
1.2
99.6
0.0
99.9
87.3
M2-Top2 †
Sigmoid multi-label classifier
9.32M
√
√
√
88.6
19.0
99.6
0.0
99.9
82.0
TABLE IV: Subset-wise exact set accuracy (%) on CS-FLEURS. Seen and Unseen denote whether each language pair appears in CS training data. † marks oracle Top-2. Other notation follows Table III .
ID
Model family
#Params
FL
CS-FL
CS-YO
Read Seen
Read Unseen
XTTS1 Seen
XTTS2 Unseen
MMS Seen
MMS Unseen
H2
Hard-target multi-class classifier
9.33M
√
√
–
0.2
0.0
99.9
99.9
99.9
38.5
H3
Hard-target multi-class classifier
9.33M
√
√
√
0.8
0.0
99.8
99.7
99.9
14.6
G1
Multi-label set generator (ours)
23.55M
√
√
–
0.0
0.0
100.0
100.0
99.9
96.8
G2
Multi-label set generator (ours)
23.55M
√
√
√
0.2
0.0
100.0
99.9
99.7
43.1
TABLE V: Two-language prediction rate (%) on CS-FLEURS. For hard-target classifiers, atomic-pair predictions count as two-language predictions; for set generators, the rate is computed from unique generated language tokens. Notation follows Table III .
Fig. 1: Threshold sensitivity of the soft-target (left) and sigmoid (right) classifiers trained on FLEURS and CS-FLEURS (S1 and M1). Curves show exact set accuracy obtained by thresholding softmax probabilities or sigmoid scores; dotted lines show the corresponding oracle Top- k accuracies. No single threshold recovers Top- k accuracy across subsets.
English proportion
H2
S1-Top2 †
M1-Top2 †
G1
All
0.18
16.46
26.16
0.07
0–10%
0.00
57.83
70.87
0.00
10–20%
0.25
28.26
43.49
0.25
20–30%
0.00
20.65
32.59
0.00
30–40%
0.39
15.69
24.71
0.00
40–50%
0.00
17.65
25.88
0.00
TABLE VI: Exact set accuracy (%) on the Bangor Miami English–Spanish subset, grouped by English proportion. † marks oracle Top-2. Bold marks row maxima.
Theta One Korea, Seoul, Republic of Korea · Kitsch Labs, Seongnam, Republic of Korea · Department of Artificial Intelligence, University of Seoul, Seoul, Republic of Korea +1