Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in principle for a typologically identifiable class of concepts, including pragmatic markers, honorifics, and diachronically stratified terms. We formalize this failure using a usage-cloud framework, representing concepts as point sets of contextualized embeddings. We define α-unalignability as the impossibility of any mapping that simultaneously preserves lexical faithfulness (centroid correspondence) and structural faithfulness (local neighborhood topology). We provide three layers of evidence. Behaviorally, we show that FLORES-200 translation failures are predicted by language family and resource class but not by script, and that LOBSTER reasoning scores vary by family. Mechanistically, we report a Representation-Intervention Gap (RIG) in a nine-model case study on Yami: the models' activations encode a regularity along which Yami groups with other low-resource and Austronesian languages, yet interventions on language-specific neurons show no demonstrated advantage over random masks: the regularity is visible but not usable by this intervention. Finally, we operationalize these findings into a multidimensional diagnostic profile: Cycle-Consistency, Pragmatic-Load Disagreement, Manifold-Curvature Mismatch, and RIG. We argue that collapsing cultural competence into a single scalar incentivizes "probabilistic flattening," and that recognizing the unalignable class is a precondition for AI that respects, rather than erases, cultural divergence. This suggests that multilingual alignment is not a single well-defined objective, but a set of mutually incompatible projections.
Figures & tables
Profile
Diagnosis
Low CC, low PLD, low MCM, low RIG
Genuinely alignable; standard translation suffices.
High CC, low PLD, low RIG
Lexical drift but pragmatic competence intact ( scaling-tractable ).
Low CC, high PLD
Pragmatic load lost. The most common failure of current benchmarks.
High MCM, high RIG
Geometry differs and the model cannot recruit the difference. The hard case.
Any high, RIG ≈0
Failure is behavioral, not architectural; better data plausibly closes it.
Table 1: Diagnostic patterns of the four-metric unalignability profile.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Concepts as usage clouds, alignment as cloud mapping (visualization of the framework introduced in § 3 ). Top : a source cloud ΦA(xıˉn kuˇ le) exhibits internal pragmatic differentiation, with four sub-regions corresponding to superior-to-subordinate, peer-to-peer, service-encounter, and ironic uses; the target cloud ΦB(good job) has no corresponding sub-structure. The alignment map τ is judged on two conditions: (M1) centroid faithfulness, ∥τ(μA)−μB∥<ε , and (M2) structural faithfulness, preservation of rankk(v) under τ . Even when τ satisfies (M1) by aligning centroids, (M2) is necessarily violated when the target lacks the source’s pragmatic dimensionality. Bottom : the three structural routes to a large infimum in Definition 3. Route A (pragmatic-load asymmetry): many sub-registers in the source collapse to one in the target. Route B (diachronic non-stationarity): the cloud drifts across time windows t1→t2 , so no synchronic gloss can simultaneously match both. Route C (categorial non-correspondence): the target cloud is a projection of the source onto a lower-dimensional subspace.
Figure 2: Illustrative example with synthetic clouds: the Procrustes residual rProc (Metric 3, Manifold-Curvature Mismatch) for a structured source cloud and a flat target cloud. Left : the raw usage clouds Φcmn(xıˉn kuˇ le) (blue), partitioned into distinct pragmatic regions (superior-to-subordinate, peer-to-peer, service encounter), and Φeng(good job) (red). Right : the clouds after optimal Procrustes alignment (translation, rotation, and scaling); gray lines connect corresponding points. rProc is the sum of squared distances between corresponding points after alignment. Even the best linear map cannot stretch the flat target cloud over the sub-clusters of the source cloud, so rProc remains high: the internal rank structure and local neighborhood densities are incompatible, an (M2) violation.
Family
Sizes
Architecture
HuggingFace identifier
Qwen3
4B / 14B / 32B
Decoder-only
Qwen/Qwen3-{4B,14B,32B}
Gemma 3 PT
4B / 12B / 27B
Decoder-only
google/gemma-3-{4b,12b,27b}-pt
NLLB-200
3.3B
Enc–dec
facebook/nllb-200-3.3B ∗
SeaLLMs-v3
7B
Decoder-only
SeaLLMs/SeaLLMs-v3-7B
SEA-LION-v4
27B
Decoder-only
aisingapore/Gemma-SEA-LION-v4-27B
Appendix
Table 2: Models in the case study. ∗ Extended with a tao_Latn language token before fine-tuning.
ID
Condition
LAPE target
LAPE contrast
Fine-tune pair
Exp 0
Neg. Ctrl
{rus, vie, kor}
{cmn, arb, fil, eng}
Yami–Chinese
Exp 1
Yami
{tao}
{cmn, kor, rus, vie, arb}
Yami–Chinese
Exp 2a
Filipino
{fil}
{cmn, kor, rus, arb, vie, eng}
Yami–Filipino
Exp 2b
Pan-AN
{fil, ceb, ind, mri, plt}
{cmn, kor, rus, arb, vie, eng}
Yami–Filipino
Exp 3
English
{eng}
{cmn, kor, rus, fil}
Yami–English
Appendix
Table 3: Five neuron-sourcing conditions; LAPE selects the listed target-set neurons (relative to the listed contrast set) via Language Activation Probability Entropy. Each condition is fine-tuned on the listed parallel pair under three modes: Standard LoRA (no mask), LAPE (gradient-masked LoRA on the LAPE-selected neurons), and Random Mask (matched-sparsity LoRA on a random neuron subset; absent for Exp 0). The common evaluation benchmark is Yami ↔ Chinese for all conditions.
Model
Condition
Train Loss
Val. Loss
Steps
Time (h)
Qwen3-4B
Standard LoRA
0.628
2.048
5369
1.3
Standard LoRA (Yami–Chinese)
0.491
1.947
5369
1.3
Exp 0 (Neg. Ctrl) [LAPE]
0.981
2.061
8437
2.0
Exp 1 (Yami) [LAPE]
1.076
2.040
7670
1.8
Exp 1 (Yami) [Random]
0.977
2.071
8437
2.0
Exp 2a (Filipino) [LAPE]
1.347
2.072
7670
2.2
Appendix
Table 4: Training summary across models and experimental conditions. Train/validation loss are final values; time is wall-clock hours. The Standard LoRA row averages the Yami–Chinese, Yami–English, and Yami–Filipino baselines; Standard LoRA (Yami–Chinese) is the matched baseline for Exp 1. LAPE/Random entries are single runs.
Comparison
Metric
Δ
95% stab. int.
dz
Agree
Exp 0 LAPE vs. Standard LoRA (cmn)
BLEU
− 4.48
[ − 5.86, − 3.06]
− 1.96
9/9
chrF++
− 5.97
[ − 7.76, − 4.14]
− 2.04
9/9
GLEU
− 3.76
[ − 4.96, − 2.51]
− 1.87
9/9
Exp 1 LAPE vs. Standard LoRA (cmn)
BLEU
− 3.75
[ − 4.65, − 2.90]
− 2.64
9/9
chrF++
− 4.69
[ − 5.49, − 3.83]
− 3.48
9/9
GLEU
− 2.84
[ − 3.53, − 2.17]
− 2.57
9/9
Appendix
Table 5: Main model-level paired-analysis summary for the 16 pre-specified comparisons on the common Yami ↔ Chinese benchmark. Δ : model-level mean difference A−B (BLEU and GLEU on [0,100] scale; chrF++ native). 95% stability interval: percentile interval over resampled model-level differences. Agree: number of trained models whose difference shares the sign of the overall estimate. dz : paired-samples Cohen’s dz (mean paired difference divided by SD of paired differences).
Comparison
BLEU
chrF++
GLEU
Exp 0 LAPE vs. Standard LoRA (cmn)
9/9
9/9
9/9
Exp 1 LAPE vs. Standard LoRA (cmn)
9/9
9/9
9/9
Exp 2a LAPE vs. Standard LoRA (fil)
9/9
9/9
9/9
Exp 2b LAPE vs. Standard LoRA (fil)
9/9
9/9
9/9
Exp 3 LAPE vs. Standard LoRA (eng)
8/9
8/9
8/9
Exp 1 LAPE vs. Random Mask
6/9
6/9
7/9
Appendix
Table 6: Sign agreement across the three metrics: number of the nine models whose model-level difference shares the sign of the overall estimate.
Figure 3: Forest plot of model-level mean paired BLEU differences (system A − system B) with 95% percentile stability intervals for all 16 comparisons on the Yami ↔ Chinese common benchmark. Labels give the number of the nine models sharing the sign of the estimate. Comparisons are grouped by family (LAPE-vs-Standard, LAPE-vs-Random, cross-condition, negative-control validation). The LAPE-vs-Random intervals all include zero: no demonstrated LAPE advantage, which is not evidence of equivalence.
Table 7: Per-model Yami nearest neighbors. Cosine similarity computed per-layer between activation-probability profiles, then averaged across layers. The last column flags whether every Austronesian language (Cebuano, Filipino, Indonesian, Maori, Plateau Malagasy) is closer to Yami than every non-Austronesian language tested. The encoder–decoder NLLB-200 is the partial exception: Filipino, Cebuano, and Plateau Malagasy remain Yami’s nearest neighbors, but English and Korean rank above Maori and Indonesian.
Figure 4: Pairwise activation cosine similarity heatmap (Qwen3-4B; per-layer cosine, mean across layers). Yami’s highest similarities are to the Austronesian languages (Plateau Malagasy, Cebuano, Filipino, Maori, Indonesian); Mandarin, English, and the contrast set (Russian/Vietnamese/Korean/Arabic) are distal.
Figure 5: Left : hierarchical clustering (average linkage) of the sixteen non-Yami evaluation languages and Yami ( tao ) by activation distance ( 1− cosine similarity of per-layer activation profiles, averaged across layers), averaged across all nine models. The low-resource languages, including most of the Austronesian set (green labels) and the low-resource non-Austronesian controls, group with Yami, while the higher-resource languages (among them Austronesian Indonesian, ind ) form a separate cluster. Right : mean Chinese → Yami BLEU for each LAPE condition, averaged across models. The clustering is not matched by a corresponding translation gain from LAPE targeting: the Filipino and Pan-AN conditions fall well below the negative control, and direct Yami targeting (Exp 1) exceeds it only slightly; the gap holds whether the clustering reflects Austronesian genealogy, corpus-resource scale, or a mixture of the two [ Lian, 2026 ] .
Neg. Ctrl
Yami
Filipino
Pan-AN
English
Neg. Ctrl
1.000
0.000
0.007
0.009
0.003
Yami
0.000
1.000
0.152
0.128
0.013
Filipino
0.007
0.152
1.000
0.336
0.020
Pan-AN
0.009
0.128
0.336
1.000
0.018
English
0.003
0.013
0.020
0.018
1.000
Appendix
Table 8: Pairwise Jaccard similarity between experimental masks, averaged across all 9 models. Values close to 0 indicate distinct neuron populations; values close to 1 indicate near-identical masks.
Figure 6: Per-(model, mask-pair) Jaccard heatmap. The (Yami, Neg. Ctrl) cell is uniformly near zero across all nine models.
Mask
Mean Jaccard
Mean shared neurons
Korean
0.1088
138.2
Vietnamese
0.0783
102.0
Russian
0.0233
26.9
Neg. Ctrl (union)
0.0100
51.0
English
0.0014
7.3
Filipino
0.0003
3.8
Appendix
Table 9: Overlap between selected masks and the model’s LAPE-identified Mandarin mask, averaged across the nine models. This tests whether the proxy-language masks may be harming the common Yami ↔ Chinese benchmark by masking Mandarin-associated neurons. The proxy-language masks overlap it far less than the negative-control union does. Korean, Vietnamese, and Russian are separate single-language masks shown for reference.
Model
Top-1 default prediction
Count
Unique preds
Empty preds
Qwen3-4B
“small fish”
125
2652
0
Qwen3-14B
“hair”
137
2824
0
Qwen3-32B
“small boat”
270
2811
0
Gemma-3-4B
“very difficult”
57
2751
0
Gemma-3-12B
“hunting”
34
2861
0
Gemma-3-27B
“catch by hand”
45
2896
0
Appendix
Table 10: The Yami → Chinese collapse defaults to a small Yami-domain Chinese vocabulary (English glosses shown), and the default token differs by model. Filipino-LAPE condition (Exp 2a), all 3,566 test items per model. NLLB differs qualitatively: it emits whitespace rather than a default token.
Model
Long ( > 50 wc)
Max wc
CJK leak
Top token in long preds
( × )
Qwen3-4B
184
256
23
ya
2810
Qwen3-14B
121
191
1
a
2003
Qwen3-32B
29
253
2
mo
403
Gemma-3-4B
70
254
73
a
2080
Gemma-3-12B
55
256
23
am,
913
Gemma-3-27B
7
170
20
a
197
Appendix
Table 11: Pan-AN (Exp 2b) Chinese → Yami degeneration is bimodal: some models produce long Yami particle streams (no lexical content); others leak Chinese characters into the Yami output. SEA-LION-v4 leaks the most CJK despite SEA-language pretraining; Gemma-3-27B is the most stable.
Condition
Empty preds / total
Rate
Exp 1 (Yami), Random Mask
0 / 3566
0.0%
Standard LoRA, Yami–Chinese
0 / 3566
0.0%
Exp 0 (Neg. Ctrl), LAPE
0 / 3566
0.0%
Exp 1 (Yami), LAPE
0 / 3566
0.0%
Standard LoRA, Yami–English
4 / 3566
0.1%
Base model (no fine-tuning)
9 / 3566
0.3%
Appendix
Table 12: NLLB Yami → Chinese empty-prediction rate by condition (3,566 items per condition). Empty outputs are concentrated in the cross-pair conditions under sparse masking (LAPE or Random); the matched-pair and Standard-LoRA conditions stay at or below 1%.