During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.
Figures & tables
Model
Switch Rate
Switch Point
Baseline
aa ✗
aaa ✗
Static
Fixed
Random
Curriculum
Gradual
Random
POS-Static
Fixed
POS-constrained
POS-Curriculum
Gradual
POS-constrained
Table 1: Overview of code-switching variants.
eng–nld
eng–zho
Eval.
Model
Δ eng
Δ nld
Δ eng
Δ zho
ZS
Static
-0.08
+0.19
-0.52
-4.14
Curr.
-0.32
-0.19
-0.20
-3.99
POS-Static
+0.38
-0.30
-0.59
-3.66
POS-Curr.
-0.54
-0.46
-0.68
-3.79
FT
Static
+2.11
+5.23
+6.09
+0.58
Table 2: Macro-averaged performance changes relative to the bilingual baseline ( Δ in percentage points). The results are averaged across three random seeds. ZS: zero-shot; FT: fine-tuning. Best improvements within each language pair and evaluation setting are in bold .
Figure 1: Early learning curve of grammatical knowledge for English–Dutch (top row) and English–Chinese (bottom row) bilingual models. Left column: English BLiMP; right column: target-language BLiMP (Dutch BLiMP for eng–nld, Chinese BLiMP for eng–zho). Each panel shows BLiMP accuracy as a function of words seen during training. The code-switching line represents the mean across four variants. Shaded bands denote standard deviation across random seeds.
Figure 2: Similarity of the global structure of sentence representations within a language in bilingual models measured by centered kernel alignment (CKA). We compare the Baseline model against the mean of code-switching variants ( CS-mean ) for the English–Dutch ( eng – nld ) models in the top row and the English–Chinese ( eng – zho ) models in the bottom row. Darker cell colors indicate higher representational similarity.
Figure 3: Cross-lingual embedding alignment across models for English–Dutch (eng–nld; top row) and English–Chinese (eng–zho; bottom row). The left column displays the mean cosine similarity for aligned parallel sentences, while the right column shows lexical alignment for curated translation pairs (filled bars) versus randomly matched word pairs (outlined bars).
Figure 4: Relationship between orthographic similarity and word embedding cosine similarity for English–Dutch cognates ( n=5,285 pairs). Contours indicate kernel density estimates. Dotted connectors link the same cognate pair between the Baseline and POS-Static setups, illustrating shifts in embedding similarity at equivalent levels of orthographic overlap (e.g., answer/antwoord , line/lijn ), where ideal cross-lingual representations maintain high cosine similarity regardless of distance. Pearson correlations are r=0.75 for Baseline and r=0.72 for POS-Static.
Figure 5: Lexical-level Representational Similarity Analysis (RSA) heatmaps for English—Dutch ( eng – nld ) and English–Chinese ( eng – zho ). Panels compare cross-model representational similarity between the Baseline model and the average of code-switching variants ( CS-mean ) over paired-token embeddings. Darker cell colors indicate higher representational similarity.
Language pair
Model
eng → X
X → eng
eng–nld
Baseline
13.9
14.9
Static
30.5
32.0
Curr.
20.2
22.1
POS-Static
21.5
22.4
POS-Curr.
17.2
19.4
eng–zho
Baseline
0 5.6
0 6.9
Table 3: Sentence-level translation retrieval measured by P@1 (%), showing the percentage of the correct translation that appears among the top- 1 retrieved sentences. The results are averaged across three random seeds. The P@1 for random retrieval is 0.1%.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Languages
Release
Aligned Sentence Pairs
eng--nld
v2018
325,874
eng--zho
v2024
504,020
Appendix
Table 4: Parallel Corpus Configurations for Bilingual Lexicon Extraction
Parameter
Value
Model architecture
12-layer GPT-2 with hidden size 768 and 12 attention heads.
Tokenizer
16,384-token shared bilingual BPE vocabulary.
Optimizer
AdamW.
Learning Rate
5×10−5 . Cosine decay with 1% linear warmup.
Batch size
Sequence length
512 tokens.
Appendix
Table 5: Training hyperparameters and configuration.
Figure 6: Relationship between orthographic similarity and word embedding cosine similarity for English–Dutch cognates ( n=5,285 pairs). Lower Pearson correlation indicates better cross-lingual alignment, especially for cognates with less similar surface forms.
Table 6: Zero-shot evaluation benchmarks grouped by evaluation dimension. 1 Choshen et al. (2026) ; 2 Jumelet et al. (2026b) ; 3 Suijkerbuijk et al. (2025) ; 4 Liu et al. (2025) ; 5 Han et al. (2026) ; 6 He et al. (2025) ;