Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders
Organizations: Indian Institute of Science Education and Research, Bhopal, India · Microsoft Corporation
Abstract
Multilingual encoder-based language models are widely used for code-mixed analysis, yet their internal representations of code-mixed inputs -- and their relationship to the constituent languages -- remain poorly understood. Using Hindi-English as a case study, we construct a unified trilingual corpus of parallel English, Hindi (Devanagari), and Romanized code-mixed sentences. We then probe cross-lingual representation alignment in standard multilingual encoders and their code-mix-adapted variants using CKA, token-level saliency, and entropy-based uncertainty analysis. We find that while standard models align English and Hindi well, code-mixed inputs remain loosely connected to either language -- and that continued pre-training on code-mixed data improves English-code-mixed alignment at the cost of English-Hindi alignment. Interpretability analyses further reveal a clear asymmetry: models process code-mixed text through an English-dominant semantic subspace, while native-script Hindi provides complementary signals that reduce representational uncertainty. Motivated by these findings, we introduce a trilingual post-training alignment objective that brings code-mixed representations closer to both constituent languages simultaneously, yielding more balanced cross-lingual alignment and downstream gains on sentiment analysis and hate speech detection -- showing that grounding code-mixed representations in their constituent languages meaningfully helps cross-lingual understanding. Code is available at: https://github.com/debajyotimaz/tri_align_EMNLP_2026.
Figures & tables
| Dataset | # Samples | Languages |
| CM-En parallel Dhar et al. (2018) | 6,096 | CM, EN |
| PHINC Srivastava and Singh (2020) | 13,738 | CM, EN |
| LINCE 2021 1 | 8060 | CM, EN, HI |
| Model | Version |
|---|---|
| mBERT | bert-base-multilingual-cased |
| Hing-mBERT | l3cube-pune/hing-bert |
| Hing-mBERT-Mixed | l3cube-pune/hing-bert-mixed |
| XLM-R | xlm-roberta-base |
| Hing-RoBERTa | l3cube-pune/hing-roberta |
| Hing-RoBERTa-Mixed | l3cube-pune/hing-roberta-mixed |
| EN CM | EN HI | HI CM | CLAS | ||||
| Model | |||||||
| Last Layer Retrieval | |||||||
| mBERT Family | |||||||
| mBERT | 54.75 | 42.90 | 74.87 | 33.80 | 26.30 | 39.05 | 14.20 |
| mBERT Trilingual | 69.43 | 63.49 | 71.52 | 73.23 | 54.49 | 49.69 | 50.97 |
| Hing-mBERT | 52.78 | 71.86 | 32.30 | 33.43 | 24.70 | 28.99 | 17.01 |
| Model | Sentiment | Hate |
|---|---|---|
| mBERT | 0.4012 | 0.5220 |
| mBERT-Tri | 0.4898 | 0.5953 |
| XLM-R | 0.5720 | 0.6307 |
| XLM-R-Tri | 0.5986 | 0.6701 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric | Score |
|---|---|
| Average BLEU | 0.2527 |
| Average SacreBLEU | 38.73 |
| Average COMET | 0.8347 |
| EN-CM | CM-EN | EN-HI | HI-EN | HI-CM | CM-HI | CLAS | |
| Last layer | |||||||
| mBERT (Google Translate) | 54.75 | 42.90 | 74.87 | 33.80 | 26.30 | 39.05 | 14.20 |
| XLM-R (Google Translate) | 50.56 | 50.67 | 58.81 | 54.33 | 38.37 | 40.46 | 39.53 |
| mBERT (IndicTrans2) | 55.37 | 42.98 | 72.76 | 32.46 | 25.35 | 37.79 | 13.53 |
| XLM-R (IndicTrans2) | 50.15 | 50.55 | 58.34 | 54.01 | 38.17 | 40.44 | 39.28 |
| Mean layer | |||||||
| Model | Version |
|---|---|
| IndicBERT-v2 | ai4bharat/IndicBERTv2-MLM-only |
| Qwen3-Embedding | Qwen/Qwen3-Embedding-0.6B |
| Model | Last Layer | Mean |
|---|---|---|
| IndicBERTv2 | 33.59 | 34.14 |
| IndicBERTv2 Tri | 51.47 | 37.27 |
| Qwen3-Embedding | 22.08 | 50.25 |
| Qwen3-Embedding Tri | 79.59 | 67.70 |
| Task | Model | Train CM | Train EN | Train HI |
|---|---|---|---|---|
| Hate Speech Detection | ||||
| IndicBERTv2 | 0.6340 | 0.6097 | 0.6142 | |
| IndicBERTv2 Tri | 0.6576 | 0.6438 | 0.6496 | |
| Sentiment Analysis | ||||
| IndicBERTv2 | 0.6346 | 0.4932 | 0.6063 | |
| IndicBERTv2 Tri | 0.6393 | 0.5014 | 0.6154 | |
| Model Family | Training Objective | Last Layer | Mean |
|---|---|---|---|
| CLAS | |||
| mBERT Family | |||
| mBERT | Standard | 14.20 | 66.36 |
| mBERT | Tri | 50.97 | 71.02 |
| mBERT | InfoNCE | 86.78 | 70.07 |
| mBERT | AnInfoNCE | 83.79 | 73.20 |
| Model | Train CM | Train EN | Train HI |
|---|---|---|---|
| mBERT | 0.5542 | 0.4805 | 0.5314 |
| mBERT Tri | 0.6194 | 0.6119 | 0.5545 |
| mBERT InfoNCE | 0.6207 | 0.5627 | 0.5903 |
| mBERT AnInfoNCE | 0.5971 | 0.5700 | 0.6100 |
| XLM-R | 0.6217 | 0.6326 | 0.6379 |
| XLM-R-Tri | 0.6794 | 0.6732 | 0.6577 |
| Model | Train CM | Train EN | Train HI |
|---|---|---|---|
| mBERT | 0.3283 | 0.3674 | 0.5079 |
| mBERT Tri | 0.4464 | 0.4710 | 0.5519 |
| mBERT InfoNCE | 0.5433 | 0.4657 | 0.5760 |
| mBERT AnInfoNCE | 0.5553 | 0.4562 | 0.5550 |
| XLM-R | 0.6304 | 0.4686 | 0.6170 |
| XLM-R-Tri | 0.6618 | 0.5022 | 0.6318 |
| EN CM | EN HI | HI CM | CLAS | ||||
| Model | |||||||
| Last Layer Retrieval | |||||||
| mBERT Family | |||||||
| mBERT | 54.75 | 42.90 | 74.87 | 33.80 | 26.30 | 39.05 | 14.20 |
| mBERT Trilingual | 69.43 | 63.49 | 71.52 | 73.23 | 54.49 | 49.69 | 50.97 |
| –w/o Alignment | 52.90 | 47.49 | 66.83 | 30.40 | 22.91 | 36.55 | 15.06 |
| EN CM | EN HI | HI CM | CLAS | ||||
| Model | |||||||
| Last Layer Retrieval | |||||||
| mBERT Family | |||||||
| mBERT | 36.46 | 28.11 | 52.07 | 16.41 | 11.66 | 19.06 | 1.68 |
| mBERT Trilingual | 51.73 | 46.61 | 57.66 | 60.80 | 36.02 | 31.57 | 32.70 |
| Hing-mBERT | 35.11 | 46.36 | 26.85 | 14.70 | 12.03 | 10.72 | 3.82 |
| Model Family | Model | PPL (CM) |
|---|---|---|
| mBERT | mBERT | 462.41 |
| mBERT Trilingual | 119.28 | |
| Hing-mBERT | 10.98 | |
| Hing-mBERT Trilingual | 8.08 | |
| Hing-mBERT-Mixed | 10.40 | |
| Hing-mBERT-Mixed Trilingual | 8.37 |
| Label | Train | Val | Test | Total |
| Sentiment Patwa et al. (2020) | ||||
| Neutral | 5,426 | 1,124 | 1,071 | 7,621 |
| Positive | 4,853 | 979 | 979 | 6,811 |
| Negative | 4,232 | 887 | 876 | 5,995 |
| Total | 14,511 | 2,990 | 2,926 | 20,427 |
| Hate Speech Bohra et al. (2018) | ||||
| Train | Model | Test Set | Consistency | ||
|---|---|---|---|---|---|
| CM | EN | HI | |||
| CM | mBERT | 0.5192 | 0.5672 | 0.3095 | 0.3283 |
| mBERT Trilingual | 0.6290 | 0.6420 | 0.4281 | 0.4464 | |
| –w/o Alignment | 0.6537 | 0.6177 | 0.3538 | 0.3780 | |
| EN | mBERT | 0.4196 | 0.7323 | 0.4604 | 0.3674 |
| mBERT Trilingual | 0.5333 | 0.7277 | 0.5103 | 0.4710 | |
| Train | Model | Test Set | Consistency | ||
|---|---|---|---|---|---|
| CM | EN | HI | |||
| CM | Hing-mBERT | 0.6916 | 0.7141 | 0.3221 | 0.3558 |
| Hing-mBERT Trilingual | 0.6722 | 0.6895 | 0.3784 | 0.4052 | |
| Hing-mBERT Mixed | 0.7241 | 0.7122 | 0.6812 | 0.6837 | |
| Hing-mBERT Mixed Trilingual | 0.7139 | 0.7110 | 0.6900 | 0.6919 | |
| EN | Hing-mBERT | 0.5613 | 0.7311 | 0.2748 | 0.2918 |
| Train | Model | Test Set | Consistency | ||
|---|---|---|---|---|---|
| CM | EN | HI | |||
| CM | mBERT | 0.6706 | 0.6194 | 0.5373 | 0.5542 |
| mBERT Trilingual | 0.6631 | 0.6188 | 0.6427 | 0.6194 | |
| –w/o Alignment | 0.6847 | 0.6575 | 0.5937 | 0.6072 | |
| EN | mBERT | 0.5777 | 0.6294 | 0.4732 | 0.4805 |
| mBERT Trilingual | 0.6391 | 0.6362 | 0.6094 | 0.6119 | |
| Train | Model | Test Set | Consistency | ||
|---|---|---|---|---|---|
| CM | EN | HI | |||
| CM | Hing-mBERT | 0.6981 | 0.6748 | 0.3890 | 0.4468 |
| Hing-mBERT Trilingual | 0.6556 | 0.6212 | 0.5136 | 0.5363 | |
| Hing-mBERT Mixed | 0.6961 | 0.6660 | 0.6109 | 0.6224 | |
| Hing-mBERT Mixed Trilingual | 0.6915 | 0.6692 | 0.6475 | 0.6514 | |
| EN | Hing-mBERT | 0.6970 | 0.6883 | 0.3890 | 0.4482 |