Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation
Organizations: National Taiwan University
Abstract
Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in principle for a typologically identifiable class of concepts, including pragmatic markers, honorifics, and diachronically stratified terms. We formalize this failure using a usage-cloud framework, representing concepts as point sets of contextualized embeddings. We define -unalignability as the impossibility of any mapping that simultaneously preserves lexical faithfulness (centroid correspondence) and structural faithfulness (local neighborhood topology). We provide three layers of evidence. Behaviorally, we show that FLORES-200 translation failures are predicted by language family and resource class but not by script, and that LOBSTER reasoning scores vary by family. Mechanistically, we report a Representation-Intervention Gap (RIG) in a nine-model case study on Yami: the models' activations encode a regularity along which Yami groups with other low-resource and Austronesian languages, yet interventions on language-specific neurons show no demonstrated advantage over random masks: the regularity is visible but not usable by this intervention. Finally, we operationalize these findings into a multidimensional diagnostic profile: Cycle-Consistency, Pragmatic-Load Disagreement, Manifold-Curvature Mismatch, and RIG. We argue that collapsing cultural competence into a single scalar incentivizes "probabilistic flattening," and that recognizing the unalignable class is a precondition for AI that respects, rather than erases, cultural divergence. This suggests that multilingual alignment is not a single well-defined objective, but a set of mutually incompatible projections.
Figures & tables
| Profile | Diagnosis |
|---|---|
| Low CC, low PLD, low MCM, low RIG | Genuinely alignable; standard translation suffices. |
| High CC, low PLD, low RIG | Lexical drift but pragmatic competence intact ( scaling-tractable ). |
| Low CC, high PLD | Pragmatic load lost. The most common failure of current benchmarks. |
| High MCM, high RIG | Geometry differs and the model cannot recruit the difference. The hard case. |
| Any high, RIG | Failure is behavioral, not architectural; better data plausibly closes it. |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Sizes | Architecture | HuggingFace identifier |
|---|---|---|---|
| Qwen3 | 4B / 14B / 32B | Decoder-only | Qwen/Qwen3-{4B,14B,32B} |
| Gemma 3 PT | 4B / 12B / 27B | Decoder-only | google/gemma-3-{4b,12b,27b}-pt |
| NLLB-200 | 3.3B | Enc–dec | facebook/nllb-200-3.3B ∗ |
| SeaLLMs-v3 | 7B | Decoder-only | SeaLLMs/SeaLLMs-v3-7B |
| SEA-LION-v4 | 27B | Decoder-only | aisingapore/Gemma-SEA-LION-v4-27B |
| ID | Condition | LAPE target | LAPE contrast | Fine-tune pair |
|---|---|---|---|---|
| Exp 0 | Neg. Ctrl | {rus, vie, kor} | {cmn, arb, fil, eng} | Yami–Chinese |
| Exp 1 | Yami | {tao} | {cmn, kor, rus, vie, arb} | Yami–Chinese |
| Exp 2a | Filipino | {fil} | {cmn, kor, rus, arb, vie, eng} | Yami–Filipino |
| Exp 2b | Pan-AN | {fil, ceb, ind, mri, plt} | {cmn, kor, rus, arb, vie, eng} | Yami–Filipino |
| Exp 3 | English | {eng} | {cmn, kor, rus, fil} | Yami–English |
| Model | Condition | Train Loss | Val. Loss | Steps | Time (h) |
|---|---|---|---|---|---|
| Qwen3-4B | Standard LoRA | 0.628 | 2.048 | 5369 | 1.3 |
| Standard LoRA (Yami–Chinese) | 0.491 | 1.947 | 5369 | 1.3 | |
| Exp 0 (Neg. Ctrl) [LAPE] | 0.981 | 2.061 | 8437 | 2.0 | |
| Exp 1 (Yami) [LAPE] | 1.076 | 2.040 | 7670 | 1.8 | |
| Exp 1 (Yami) [Random] | 0.977 | 2.071 | 8437 | 2.0 | |
| Exp 2a (Filipino) [LAPE] | 1.347 | 2.072 | 7670 | 2.2 |
| Comparison | Metric | 95% stab. int. | Agree | ||
|---|---|---|---|---|---|
| Exp 0 LAPE vs. Standard LoRA (cmn) | BLEU | 4.48 | [ 5.86, 3.06] | 1.96 | 9/9 |
| chrF++ | 5.97 | [ 7.76, 4.14] | 2.04 | 9/9 | |
| GLEU | 3.76 | [ 4.96, 2.51] | 1.87 | 9/9 | |
| Exp 1 LAPE vs. Standard LoRA (cmn) | BLEU | 3.75 | [ 4.65, 2.90] | 2.64 | 9/9 |
| chrF++ | 4.69 | [ 5.49, 3.83] | 3.48 | 9/9 | |
| GLEU | 2.84 | [ 3.53, 2.17] | 2.57 | 9/9 |
| Comparison | BLEU | chrF++ | GLEU |
| Exp 0 LAPE vs. Standard LoRA (cmn) | 9/9 | 9/9 | 9/9 |
| Exp 1 LAPE vs. Standard LoRA (cmn) | 9/9 | 9/9 | 9/9 |
| Exp 2a LAPE vs. Standard LoRA (fil) | 9/9 | 9/9 | 9/9 |
| Exp 2b LAPE vs. Standard LoRA (fil) | 9/9 | 9/9 | 9/9 |
| Exp 3 LAPE vs. Standard LoRA (eng) | 8/9 | 8/9 | 8/9 |
| Exp 1 LAPE vs. Random Mask | 6/9 | 6/9 | 7/9 |
| Model | Top-5 nearest neighbors (cosine) | Closest non-AN | All-AN tighter? |
|---|---|---|---|
| Qwen3-4B | Plateau Malagasy (0.890), Cebuano (0.890), Filipino (0.888), Maori (0.856), Indonesian (0.821) | Arabic (0.797) | yes |
| Qwen3-14B | Filipino (0.867), Plateau Malagasy (0.866), Cebuano (0.864), Maori (0.840), Indonesian (0.797) | Arabic (0.778) | yes |
| Qwen3-32B | Filipino (0.880), Cebuano (0.876), Plateau Malagasy (0.864), Maori (0.840), Indonesian (0.808) | Arabic (0.786) | yes |
| Gemma-3-4B | Cebuano (0.789), Filipino (0.785), Plateau Malagasy (0.764), Maori (0.761), Indonesian (0.726) | Mandarin (0.707) | yes |
| Gemma-3-12B | Cebuano (0.807), Filipino (0.801), Plateau Malagasy (0.774), Maori (0.773), Indonesian (0.722) | Mandarin (0.692) | yes |
| Gemma-3-27B | Cebuano (0.777), Filipino (0.774), Plateau Malagasy (0.764), Maori (0.748), Indonesian (0.697) | Mandarin (0.681) | yes |
| Neg. Ctrl | Yami | Filipino | Pan-AN | English | |
|---|---|---|---|---|---|
| Neg. Ctrl | 1.000 | 0.000 | 0.007 | 0.009 | 0.003 |
| Yami | 0.000 | 1.000 | 0.152 | 0.128 | 0.013 |
| Filipino | 0.007 | 0.152 | 1.000 | 0.336 | 0.020 |
| Pan-AN | 0.009 | 0.128 | 0.336 | 1.000 | 0.018 |
| English | 0.003 | 0.013 | 0.020 | 0.018 | 1.000 |
| Mask | Mean Jaccard | Mean shared neurons |
|---|---|---|
| Korean | 0.1088 | 138.2 |
| Vietnamese | 0.0783 | 102.0 |
| Russian | 0.0233 | 26.9 |
| Neg. Ctrl (union) | 0.0100 | 51.0 |
| English | 0.0014 | 7.3 |
| Filipino | 0.0003 | 3.8 |
| Model | Top-1 default prediction | Count | Unique preds | Empty preds |
|---|---|---|---|---|
| Qwen3-4B | “small fish” | 125 | 2652 | 0 |
| Qwen3-14B | “hair” | 137 | 2824 | 0 |
| Qwen3-32B | “small boat” | 270 | 2811 | 0 |
| Gemma-3-4B | “very difficult” | 57 | 2751 | 0 |
| Gemma-3-12B | “hunting” | 34 | 2861 | 0 |
| Gemma-3-27B | “catch by hand” | 45 | 2896 | 0 |
| Model | Long ( 50 wc) | Max wc | CJK leak | Top token in long preds | ( ) |
|---|---|---|---|---|---|
| Qwen3-4B | 184 | 256 | 23 | ya | 2810 |
| Qwen3-14B | 121 | 191 | 1 | a | 2003 |
| Qwen3-32B | 29 | 253 | 2 | mo | 403 |
| Gemma-3-4B | 70 | 254 | 73 | a | 2080 |
| Gemma-3-12B | 55 | 256 | 23 | am, | 913 |
| Gemma-3-27B | 7 | 170 | 20 | a | 197 |
| Condition | Empty preds / total | Rate |
|---|---|---|
| Exp 1 (Yami), Random Mask | 0 / 3566 | 0.0% |
| Standard LoRA, Yami–Chinese | 0 / 3566 | 0.0% |
| Exp 0 (Neg. Ctrl), LAPE | 0 / 3566 | 0.0% |
| Exp 1 (Yami), LAPE | 0 / 3566 | 0.0% |
| Standard LoRA, Yami–English | 4 / 3566 | 0.1% |
| Base model (no fine-tuning) | 9 / 3566 | 0.3% |