During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.
Figures & tables
Model
Switch Rate
Switch Point
Baseline
aa ✗
aaa ✗
Static
Fixed
Random
Curriculum
Gradual
Random
POS-Static
Fixed
POS-constrained
POS-Curriculum
Gradual
POS-constrained
Table 1: Overview of code-switching variants.
eng–nld
eng–zho
Eval.
Model
Δ eng
Δ nld
Δ eng
Δ zho
ZS
Static
-0.08
+0.19
-0.52
-4.14
Curr.
-0.32
-0.19
-0.20
-3.99
POS-Static
+0.38
-0.30
-0.59
-3.66
POS-Curr.
-0.54
-0.46
-0.68
-3.79
FT
Static
+2.11
+5.23
+6.09
+0.58
Table 2: Macro-averaged performance changes relative to the bilingual baseline ( Δ in percentage points). The results are averaged across three random seeds. ZS: zero-shot; FT: fine-tuning. Best improvements within each language pair and evaluation setting are in bold .
Figure 1: Early learning curve of grammatical knowledge for English–Dutch (top row) and English–Chinese (bottom row) bilingual models. Left column: English BLiMP; right column: target-language BLiMP (Dutch BLiMP for eng–nld, Chinese BLiMP for eng–zho). Each panel shows BLiMP accuracy as a function of words seen during training. The code-switching line represents the mean across four variants. Shaded bands denote standard deviation across random seeds.
Figure 2: Similarity of the global structure of sentence representations within a language in bilingual models measured by centered kernel alignment (CKA). We compare the Baseline model against the mean of code-switching variants ( CS-mean ) for the English–Dutch ( eng – nld ) models in the top row and the English–Chinese ( eng – zho ) models in the bottom row. Darker cell colors indicate higher representational similarity.
Figure 3: Cross-lingual embedding alignment across models for English–Dutch (eng–nld; top row) and English–Chinese (eng–zho; bottom row). The left column displays the mean cosine similarity for aligned parallel sentences, while the right column shows lexical alignment for curated translation pairs (filled bars) versus randomly matched word pairs (outlined bars).
Figure 4: Relationship between orthographic similarity and word embedding cosine similarity for English–Dutch cognates ( n=5,285 pairs). Contours indicate kernel density estimates. Dotted connectors link the same cognate pair between the Baseline and POS-Static setups, illustrating shifts in embedding similarity at equivalent levels of orthographic overlap (e.g., answer/antwoord , line/lijn ), where ideal cross-lingual representations maintain high cosine similarity regardless of distance. Pearson correlations are r=0.75 for Baseline and r=0.72 for POS-Static.
Figure 5: Lexical-level Representational Similarity Analysis (RSA) heatmaps for English—Dutch ( eng – nld ) and English–Chinese ( eng – zho ). Panels compare cross-model representational similarity between the Baseline model and the average of code-switching variants ( CS-mean ) over paired-token embeddings. Darker cell colors indicate higher representational similarity.
Language pair
Model
eng → X
X → eng
eng–nld
Baseline
13.9
14.9
Static
30.5
32.0
Curr.
20.2
22.1
POS-Static
21.5
22.4
POS-Curr.
17.2
19.4
eng–zho
Baseline
0 5.6
0 6.9
Table 3: Sentence-level translation retrieval measured by P@1 (%), showing the percentage of the correct translation that appears among the top- 1 retrieved sentences. The results are averaged across three random seeds. The P@1 for random retrieval is 0.1%.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Languages
Release
Aligned Sentence Pairs
eng--nld
v2018
325,874
eng--zho
v2024
504,020
Appendix
Table 4: Parallel Corpus Configurations for Bilingual Lexicon Extraction
Parameter
Value
Model architecture
12-layer GPT-2 with hidden size 768 and 12 attention heads.
Tokenizer
16,384-token shared bilingual BPE vocabulary.
Optimizer
AdamW.
Learning Rate
5×10−5 . Cosine decay with 1% linear warmup.
Batch size
Sequence length
512 tokens.
Appendix
Table 5: Training hyperparameters and configuration.
Figure 6: Relationship between orthographic similarity and word embedding cosine similarity for English–Dutch cognates ( n=5,285 pairs). Lower Pearson correlation indicates better cross-lingual alignment, especially for cognates with less similar surface forms.
Table 6: Zero-shot evaluation benchmarks grouped by evaluation dimension. 1 Choshen et al. (2026) ; 2 Jumelet et al. (2026b) ; 3 Suijkerbuijk et al. (2025) ; 4 Liu et al. (2025) ; 5 Han et al. (2026) ; 6 He et al. (2025) ;
Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual transfer and multilingual alignment in large language models (LLMs). However, existing studies primarily focus on bilingual transfer between English and a target language, leaving multilingual settings involving three or more languages largely unexplored. In this work, we investigate multilingual code-switching instruction tuning across four languages: English, Japanese, Korean, and Chinese. We evaluate multilingual understanding on Belebele. Our experiments show that simple sentence-level multilingual CSD consistently improves average multilingual performance across all four languages, indicating that multilingual code-switching can be effective beyond bilingual transfer settings.
Code-switching (CSW) remains challenging for large multi-lingual ASR systems in real-world deployment. While fine-tuning on synthetic CSW data is possible, it generally degrades strong monolingual baselines. Our goal is to preserve these capabilities while extending models to handle complex code-switching, including morphological variations across languages. We propose Bayesian factorized adaptation, which learns to efficiently integrate switching-relevant knowledge into strong pretrained models without overwriting existing capabilities. Requiring only a small amount of synthetic data, our approach reduces transcription errors by 32.87% on code-switched words while improving overall WER by 5.31%, all while maintaining mono-lingual performance. Our results demonstrate that effective CSW adaptation depends more on knowledge integration than data complexity.
Enes Yavuz Ugan, Alexander Waibel
Interactive Systems Lab, Karlsruhe Institute of Technology (KIT), Germany · InterACT, Carnegie Mellon University (CMU), USA
Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks. Interestingly, while reasoning models are often trained to generate monolingual text, these models have also been observed to code-switch (i.e., mix languages). Prior works have either viewed code-switching as an undesirable error, attempted to control code-switching through modifications to input prompts or the output decoding process, or focus on narrow subsets of languages, domains, tasks, and models. We address these gaps by introducing the first linguistically and behaviorally motivated fine-tuning framework for identifying beneficial code-switched reasoning behaviors in large language models and teaching these models to code-switch more effectively for reasoning. First, we create and systematically analyze a dataset of reasoning traces from diverse models, languages, tasks, and domains to understand the types of code-switching behaviors found in existing reasoning models. Then, we develop fine-tuning interventions that teach reasoning models to code-switch based on our observations of helpful behaviors in existing models. We find that our framework can significantly increase beneficial code-switched reasoning behaviors in a data-efficient manner. Interestingly, we also find that code-switching behaviors in reasoning models can be modified by fine-tuning for tasks that do not directly demonstrate code-switching in reasoning (e.g., machine translation). Our work suggests that data-efficient interventions can instill helpful forms of code-switching behavior in reasoning models.