Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.
Figures & tables
Figure 1: Overview of CoSE-E Generation Pipeline
Language Pair
Dataset
Matrix Language
Embedded Language
# utts
words/utt
chars/utt
switches/utt
CMI
switch density
% mono ( ↓ )
ES/EN
CoSE-E
ES (64.2%)
EN (35.8%)
259
24.60
114.76
6.91
0.32
0.29
0.00
CS-FLEURS
ES (80.7%)
EN (19.2%)
1576
26.01
129.33
5.90
0.16
0.24
3.24
SWITCHLINGUA
ES (64.5%)
EN (35.5%)
8110
54.95
281.99
6.71
0.32
0.13
0.95
FR/EN
CoSE-E
FR (57.2%)
EN (42.8%)
298
21.56
98.54
4.95
0.33
0.24
0.00
CoSE-E (CA)
FR-CA (73.5%)
EN (26.5%)
188
20.82
98.53
5.74
0.26
0.29
0.00
CS-FLEURS
FR (72.2%)
EN (27.7%)
1315
25.46
129.82
6.17
0.19
0.25
3.42
Table 1: Linguistic characteristics of contemporary code-switching benchmarks . CoSE-E demonstrates consistent code-switching patterns across languages (CMI 0.26–0.33, switch density 0.13–0.29) with no monolingual utterances, unlike baseline corpora which contain significant monolingual content (3.2–67.3%). Column abbreviations: # utts = utterance count ; words/utt = average words per utterance ; chars/utt = average characters per utterance ; switches/utt = code-switches per utterance ; CMI = code-mixing index ; switch density = proportion of switching frames ; % mono = percentage of monolingual content . All values reported to two decimal places.
ES/EN
FR/EN
FR-CA/EN
DE/EN
ZH/EN
Model
WER
SWER
AER
WER
SWER
AER
WER
SWER
AER
WER
SWER
AER
WER
CER
SWER
AER
Scribe-V2
.022
.004
.033
.031
.006
.051
.041
.005
.030
.027
.002
.021
.073
.031
.006
.042
Gemini-3-Flash
.028
.005
.031
.040
.009
.054
.055
.008
.043
.046
.003
.023
.090
.041
.009
.059
AssemblyAI
.029
.004
.033
.039
.009
.062
.052
.006
.034
.048
.003
.023
.093
.046
.010
.057
Qwen3-Omni
.042
.004
.027
.061
.010
.055
.071
.009
.062
.063
.004
.033
.040
.039
.006
.041
Voxtral
.049
.005
.036
.060
.012
.068
.059
.014
.074
.081
.003
.027
–
–
–
–
Table 2: WER, SWER, and AER by model and language pair. ZH/EN WER is jieba word-level (mean of per-utterance WER, as for all pairs) and broadly comparable to the Latin-script pairs; ZH/EN CER is character-level and not magnitude-comparable. All ZH/EN outputs were normalized Traditional → Simplified (OpenCC) before WER/CER/SWER scoring. Best per column in bold; ties at displayed precision (3 decimals) are both bolded.
Language pair
WER
SWER
AER
EN/DE
2.62
1.38
1.50
EN/ES
1.75
2.00
2.25
EN/FR-CA
3.62
3.13
2.88
EN/FR
2.00
3.50
3.38
Friedman χ2(3)
10.05
13.95
9.45
p
.018
.003
.024
Table 3: Mean rank of each language pair per metric (4 pairs, 8 models). Rank 1 = easiest, rank 4 = hardest; averaged across all 8 models. Bold = easiest pair per metric. Friedman χ2 and Kendall’s W summarise whether pairs differ significantly and how consistently models agree on the ordering.
Figure 2: Mean per-utterance code-switching cost , Δ=m(cs)−m(mono) ; positive values indicate code-switching is harder. Rows: WER, SWER, AER. Columns: the two monolingual baselines. Bars are grouped by model (top/mid tier) and colored by language pair; whiskers are 95% bootstrap CIs. Stars mark significance under a two-sided sign-flip permutation test, Holm-corrected within each metric × pair × baseline family ( ∗p<.05 , ∗∗p<.01 , ∗∗∗p<.001 ). Whisper is omitted; its deltas reach +2.5 (see Figure 4 ).
P(AER worsens)
Baseline
entity err.
no err.
RR
OR
English
0.44
0.054
8.3
14.0
non-English
0.39
0.052
7.6
11.8
Table 4: Entity-error propagation to task failure . RR = risk ratio; OR = odds ratio. When code-switching introduces a critical-entity error, the chance that the downstream answer worsens rises roughly eightfold, near-identically against both monolingual baselines (per-pair RR 6 – 13 , all p<10−30 ). Whisper excluded; records inner-joined by ID.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: The AER pipeline ( Pulikodan et al., 2025 ) . LLM1 generates questions from the reference, LLM2 answers from both the reference and ASR hypothesis, and LLM3 judges answer alignment.
Exp 1 (Judge)
Exp 2 (LLM2+Judge)
Flip direction
Lang.
Qs
Recs
Unan.
Non-u.
Unan.
Non-u.
Flipped
→ match
→ mis.
Flip %
EN/ES
777
259
776 (99.9%)
1
766 (98.6%)
11
3
2
1
0.4%
EN/DE
519
173
519 (100.0%)
0
515 (99.2%)
4
5
3
2
1.0%
EN/FR
894
298
893 (99.9%)
1
886 (99.1%)
8
7
4
3
0.8%
EN/FR CA
564
188
561 (99.5%)
3
559 (99.1%)
5
4
3
1
0.7%
Appendix
Table 5: AER pipeline stochasticity across language pairs. Non-unan. = questions where the 3 trials did not all agree. Flipped = records where the majority verdict changed between Experiment 1 and Experiment 2, isolating the effect of LLM2 resampling beyond judge noise.
Language
Annotator
Passed
Total
Pass Rate
AC1
DE
Annotator 1
518
519
99.8%
0.980
Annotator 2
518
519
99.8%
Annotator 3
506
519
97.5%
FR
Annotator 1
859
879
97.7%
0.978
Annotator 2
891
894
99.7%
Annotator 3
881
888
99.2%
Appendix
Table 6: Per-annotator validation pass rates and Gwet’s AC1 inter-annotator agreement for the code-switched QA validation task. Each annotator independently judged whether every generated question had an answer grounded in the reference utterance. Pass Rate gives the proportion of questions marked valid by that annotator; AC1 is computed once per language across all annotators jointly.
Language
Male Voice
Female Voice
Model
Format
English (monolingual)
Adam
Matilda
V2
24 kHz, 16-bit PCM
German
Finn
Johanna
V2
24 kHz, 16-bit PCM
French Canadian
Felix Tabarnak
Amelie
V2
24 kHz, 16-bit PCM
French
Denis Landrieu
Marine
V2
24 kHz, 16-bit PCM
Spanish
Rodrigo
Cristina Campos
V2
24 kHz, 16-bit PCM
Mandarin Chinese
Jing
Macy
V3
24 kHz, 16-bit PCM
Appendix
Table 7: TTS voices used in the benchmark. Utterances were alternated between male and female voices to achieve 50/50 gender balance. All audio was synthesized at 24 kHz, 16-bit PCM with stability=0.5, similarity_boost=0.75, style=0.2, and speaker boost enabled.
Provider
Model ID
Language-ID Setting
Decoding Parameters
AssemblyAI / Universal-3-Pro
universal-3.5-pro
language_detection: True
defaults, streaming: false
ElevenLabs / Scribe-V2
scribe_v2
language_code omitted
defaults
OpenAI / Whisper-Large-V3-Turbo
whisper-large-v3-turbo
language omitted
translate: false, defaults
Mistral / Voxtral-Small-24B
Voxtral Small 1.0 (24B)
language omitted
temperature: 0
Deepgram / Nova-3-Multilang
nova-3-multilang
language: "multi"
smart_format: true
NVIDIA / Parakeet-TDT-0.6B-V3
parakeet-tdt-0.6b-v3
N/A
defaults
Appendix
Table 8: ASR systems and decoding parameters. All models were run with no language-ID passed.
Metric
χ2(7)
N
p
Within W
WER
1464.1
918
<.001
0.23
SWER
612.2
841
<.001
0.10
AER
200.1
918
<.001
0.03
Appendix
Table 9: Omnibus and concordance statistics on the European language pairs. Friedman tests confirm that the eight models differ significantly on every metric; the reduced N for SWER reflects complete-case analysis after judge-failure drops. Within-metric Kendall’s W measures how consistently clips agree on model ordering, and declines monotonically from surface form to answer-equivalence (WER → SWER → AER), indicating that models trade places more freely as metrics move toward semantics. Cross-metric agreement across the three metric orderings is high (Cross Metric Kendall’s W=0.963 , 95% CI [0.889,0.979] ).
Condition
WER
SWER
AER
Before T → S normalisation
0.167
0.010
0.059
After T → S normalisation
0.090
0.010
0.059
Δ
−46%
—
—
Appendix
Table 10: Effect of Traditional → Simplified (OpenCC) normalisation on Gemini-3-Flash(en_zh, N=294 ). WER drops 46% once script variants are collapsed before scoring. SWER and AER are unchanged because they judge semantic equivalence, not surface form.
Figure 4: Mean per-utterance code-switching cost for Whisper-only
Language pair
WER
SWER
AER
EN/DE
2.80
1.60
1.60
EN/ES
1.20
1.60
1.80
EN/FR
2.20
4.00
4.40
EN/FR-CA
3.80
3.20
3.00
EN/ZH
5.00
4.60
4.20
Friedman χ2(4)
17.12
15.04
13.60
Appendix
Table 11: Mean rank of each language pair per metric, including ZH/EN (5 pairs, 5 models). Rank 1 = easiest, rank 5 = hardest; averaged across the 5 models with ZH/EN coverage. Bold = easiest pair per metric; SWER ties DE/EN and ES/EN. ZH/EN ranks last on every metric. WER for ZH/EN uses jieba word-level segmentation; see text for comparability caveats.
WER
SWER
AER
Model
EN
non-EN
EN
non-EN
EN
non-EN
FR/EN (10/14, 13/14, 4/14 significant)
AssemblyAI
+0.019 ***
+0.015 **
+0.008 ***
+0.005 *
+0.042 **
+0.018
Scribe-v2
+0.023 ***
+0.017
+0.005 ***
+0.003 **
+0.022
+0.018
Gemini-3-Flash
+0.027 ***
−0.001
+0.007 ***
+0.007 **
+0.022
+0.010
Qwen3-Omni
+0.031 ***
+0.029 ***
+0.010 ***
+0.008 ***
+0.028 *
+0.015
Appendix
Table 12: Per-model code-switching deltas by language pair, metric, and baseline. Each cell shows the mean per-utterance Δ=m(cs)−m(mono) ; positive values indicate code-switching raised error. Significance: * p<.05 , ** p<.01 , *** p<.001 (two-sided sign-flip permutation, Holm-corrected within each metric × pair × baseline family). Parenthesized counts in each panel header show significant cells out of total across both baselines. Whisper excluded (see Appendix D.3 ).
ES/EN
FR/EN
FR-CA/EN
DE/EN
ZH/EN
Panel 1: CMI
Whisper
1.010
0.975
0.949 *
0.969
0.950
Scribe v2
1.028 *
1.008
1.012
1.002
0.994
Nova-3
1.010
0.991
1.007
0.981
–
Parakeet
1.026
1.010
1.000
1.002
–
Voxtral-24B
1.022
0.991
1.032
1.014
–
Appendix
Table 13: Part A odds ratios from per-model, per-language-pair logistic regressions of error occurrence ( WER>0 ) on CMI, switch count, and log(nwords) . Stars denote significance: * p<.05 ; ** p<.01 ; *** p<.001 . † AssemblyAI Universal-3.5-Pro. Dashes mark unsupported language pairs.
Language pair
Utterances
nwords
Switch count
CMI
(mean ± sd, range)
(mean ± sd, range)
(mean ± sd, range)
ES/EN
259
24.6±7.7 (10–45)
6.91±3.41 (1–18)
31.7±10.5 (9.4–50.0)
FR/EN
298
21.6±6.8 (8–41)
4.95±2.83 (1–14)
33.5±10.4 (10.5–50.0)
FR-CA/EN
188
20.8±7.0 (9–42)
5.74±2.87 (1–15)
25.8±10.1 (4.0–50.0)
DE/EN
173
21.1±6.7 (9–40)
4.99±2.61 (1–15)
30.0±11.2 (5.4–50.0)
ZH/EN
294
21.2±6.3 (12–39)
2.86±1.59 (1–12)
10.3±7.9 (2.6–50.0)
Appendix
Table 14: Dataset composition per language pair. ZH/EN embeds markedly fewer switches and lower CMI than the other pairs at comparable utterance length.
Model
ES/EN
FR/EN
FR-CA/EN
DE/EN
ZH/EN
Whisper
57.1
82.2
88.3
96.5
96.3
Scribe v2
37.8
47.3
56.4
37.6
31.6
Nova-3
52.9
61.4
80.9
69.9
–
Parakeet
76.8
65.1
81.4
63.0
–
Voxtral-24B
41.3
58.4
70.2
68.8
–
Gemini-3
36.3
50.0
68.6
53.2
54.4
Appendix
Table 15: Error base rate per model × language pair (% of utterances with WER>0 ). Values near 96–97% indicate a ceiling that limits Part A’s power to detect predictor effects. † Universal-3.5-Pro. Dashes mark unsupported pairs.
Language pair
rsw,n
rsw,CMI
rn,CMI
max VIF
ES/EN
0.66
−0.04
−0.17
1.86
FR/EN
0.56
−0.06
−0.06
1.46
FR-CA/EN
0.65
0.19
−0.20
2.13
DE/EN
0.60
−0.13
−0.18
1.60
ZH/EN
0.28
0.58
−0.05
1.76
Appendix
Table 16: Predictor collinearity among switch count (sw), lognwords ( n ), and CMI. All VIFs fall well below the conventional threshold of 5. ZH/EN is the sole pair where CMI and switch count overlap substantially ( r=0.58 ).
Reference
Je peux accéder à host-2854.corp.local sur mon nouveau laptop. Par contre, I need to set up le VPN pour host-1375.corp.local —vous pouvez me walk through ca?
Table 17: Destructive entity error: every code-switched hypothesis corrupts at least one hostname (AER = 1.0), while both monolingual hypotheses recover them exactly (AER = 0.0).
Reference
…la query sur sn_safe_story …l’opération query_match sur sn_safe_story
CS hyp.
(AssemblyAI) SN_Safestory … query match … SN_Safe AER = 0.0
(French) …la requête sur sn_safe_story … query_match AER = 0.0
Appendix
Table 18: Cosmetic entity error: casing and separator changes yield several word-level entity-WER errors, but the values stay recoverable, so AER remains 0.0 across all conditions.
Entity WER
Δ vs.
Category
CS
EN
non-EN
NCS
EN
non-EN
IDs
.158
.038
.136
2,023
+ .120
+ .022
Email
.293
.183
.269
651
+ .111
+ .025
URL
.160
.050
.187
3,094
+ .110
− .027
Phone
.374
.286
.333
91
+ .088
+ .040
Address
.310
.226
.243
371
+ .084
+ .067
Appendix
Table 19: Entity WER by category. Against the English baseline, code-switching raises error on IDs, emails, URLs, and phones ( +0.09 – 0.12 ); against the non-English baseline the overall effect inverts ( −0.02 ), with only small residual increases on IDs and email. Names are unaffected in either comparison. NCS : entity word count in the code-switched condition. Date and measurement counts are not comparable across baselines due to localization and are omitted from the non-English column. Whisper excluded; micro-averaged across models and language pairs.
Entity Error
No Entity Error
Baseline
Pair
Worse
Flat
Worse
Flat
N
P EE
EN
ES/EN
93
152
87
1,441
1,773
.380
FR/EN
157
159
106
1,628
2,050
.497
FR-CA/EN
92
106
61
1,038
1,297
.465
DE/EN
38
61
38
1,049
1,186
.384
non-EN
ES/EN
70
136
63
1,520
1,789
.340
Appendix
Table 20: Per-pair entity-error propagation. P EE is the probability that AER worsens given an entity error was introduced by code-switching; the no-error base rate is ∼ 0.05 throughout. All per-pair χ2 tests yield p<10−30 . Whisper excluded; records inner-joined by ID.
Code-switching (CSW) remains challenging for large multi-lingual ASR systems in real-world deployment. While fine-tuning on synthetic CSW data is possible, it generally degrades strong monolingual baselines. Our goal is to preserve these capabilities while extending models to handle complex code-switching, including morphological variations across languages. We propose Bayesian factorized adaptation, which learns to efficiently integrate switching-relevant knowledge into strong pretrained models without overwriting existing capabilities. Requiring only a small amount of synthetic data, our approach reduces transcription errors by 32.87% on code-switched words while improving overall WER by 5.31%, all while maintaining mono-lingual performance. Our results demonstrate that effective CSW adaptation depends more on knowledge integration than data complexity.
Enes Yavuz Ugan, Alexander Waibel
Interactive Systems Lab, Karlsruhe Institute of Technology (KIT), Germany · InterACT, Carnegie Mellon University (CMU), USA
Code-switching -- the natural alternation between two languages within a single utterance -- remains one of the most challenging and under-studied conditions for automatic speech recognition (ASR). We present a benchmark evaluating five commercial ASR providers across four language pairs: Egyptian Arabic--English, Saudi Arabic (Najdi/Hijazi)--English, Persian (Farsi)--English, and German--English, comprising 300 samples per pair selected by a two-stage pipeline combining heuristic filtering with a GPT-4o and Gemini 1.5 Pro ensemble scorer, reducing LLM costs by ≈91%. We evaluate on both WER and BERTScore, showing that while both metrics agree on the ordinal ranking of systems for all Arabic and Persian pairs (τ=1.0), WER inflates the magnitude of quality gaps by approximately 3× by penalising semantically correct transliteration choices. ElevenLabs Scribe v2 achieves the lowest WER (13.2% overall) and leads on BERTScore (0.936 overall). Difficulty-stratified analysis reveals performance gaps masked by aggregate averages, and BERT embedding projections confirm semantic proximity between reference and hypothesis despite surface-level script differences. The dataset is publicly available at https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.
Sajjad Abdoli, Ghassan Al-Sumaidaee, Clayton W. Taylor +2
Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.
Yue Heng Yeo, Haoyang Li, Yizhou Peng +6
College of Computing and Data Science, Nanyang Technological University, Singapore · HLT-COE & CLSP, Johns Hopkins University, USA · Institute for Infocomm Research (I2R), A⋆STAR, Singapore +1