cs.CLMar 20, 2026

Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders

Authors: Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali, Jasabanta Patro

Organizations: Indian Institute of Science Education and Research, Bhopal, India · Microsoft Corporation

Abstract

Multilingual encoder-based language models are widely used for code-mixed analysis, yet their internal representations of code-mixed inputs -- and their relationship to the constituent languages -- remain poorly understood. Using Hindi-English as a case study, we construct a unified trilingual corpus of parallel English, Hindi (Devanagari), and Romanized code-mixed sentences. We then probe cross-lingual representation alignment in standard multilingual encoders and their code-mix-adapted variants using CKA, token-level saliency, and entropy-based uncertainty analysis. We find that while standard models align English and Hindi well, code-mixed inputs remain loosely connected to either language -- and that continued pre-training on code-mixed data improves English-code-mixed alignment at the cost of English-Hindi alignment. Interpretability analyses further reveal a clear asymmetry: models process code-mixed text through an English-dominant semantic subspace, while native-script Hindi provides complementary signals that reduce representational uncertainty. Motivated by these findings, we introduce a trilingual post-training alignment objective that brings code-mixed representations closer to both constituent languages simultaneously, yielding more balanced cross-lingual alignment and downstream gains on sentiment analysis and hate speech detection -- showing that grounding code-mixed representations in their constituent languages meaningfully helps cross-lingual understanding. Code is available at: https://github.com/debajyotimaz/tri_align_EMNLP_2026.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

    Sep 11, 2026Pruthwik Mishra, Rudra Trivedi, Avi Patel +2Code-SwitchingMultilingual

  2. Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

    Jul 28, 2026Ayushi Pandey, Pamir Gogoi, Kevin TangGrapheme-To-PhonemeIndian Languages

  3. Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

    May 28, 2026Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali +3Cross-Lingual ConsistencyIndian Languages