Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract levels are shared across languages while surface levels are modality- or language-specific. Concretely, we generate synthetic languages from probabilistic context-free grammars sharing upper-level but not lower-level production rules. In this setting the Bayes-optimal next-token predictor is belief propagation (BP); encoding its messages in successive layers yields analytical predictions that agree well with transformers trained on the same data. The framework explains why cross-lingual similarity peaks in middle layers, coexists with language-specific structure, and strengthens with language proximity, model quality and data exposure. It distinguishes similarity (shared neighborhood geometry) from alignment (shared coordinates), showing that the latter occurs when code-switched data, i.e. mixed-language sentences, are abundant enough. It further predicts that subtracting from each layer the component linearly predictable from the preceding one increases cross-lingual similarity, which we confirm in pretrained LLMs.
Figures & tables
Figure 1: (a) Multilingual Random Hierarchy Model. Upper-level production rules are shared between languages, while lower-level rules are language-specific. Zigzag line denotes this separation. Mixed sentences are generated by replacing language-specific subtrees. (b) Optimal next-token prediction via inference with belief propagation (BP). Encoding: From t observed tokens, beliefs on higher-and-higher latent variables are sequentially computed and propagated upward. The highlighted nodes denote the encoded beliefs, and the lined boxes denote the computations. Decoding: Posterior probabilities of latent variables are computed iteratively using the encoded information, toward lower-and-lower levels, ultimately leading to the probability of the (t+1) -th token. From this exact algorithm, a theoretical network is constructed, whose layerwise representations encode the steps of the BP algorithm. The computations below Lsurface are language-specific, while those above Lsurface have shared components. This leads to a maximal cross-lingual similarity near the center of theoretical networks, confirmed in trained transformers. See Sec. 4 for details. (c) Cross-lingual similarity measured with Information Imbalance (where lower Information Imbalance corresponds to higher similarities, see Appendix D ). OLMo-3-7B tested on English and Polish ( Team Olmo et al., 2026 ) , transformers trained on multilingual RHM data, and theoretical networks constructed with BP – all exhibit similar II profiles, with strongest similarity in the middle layers.
Figure 2: (a,b) Representations of a transformer trained on MRHM are geometrically similar in the middle layers when processing pairs of translated sentences. The geometric similarity is stronger for structurally similar languages, controlled by Lsurface in our setup. (c,d) This similarity also exists between representations of independent models, trained separately on languages A and B. Measurements with II and mKNN show similar trends. In all cases, the empirical trends are well-captured by theoretical networks constructed with BP.
Figure 3: (Left) Cross-lingual similarity gets stronger as the training-set size P increases. (Right) Our theory captures this behaviour with the imperfect-BP algorithm, which only knows the RHM rules up to level ℓ∗ .
Figure 4: Cross-lingual geometric similarity with low-resource languages. (a) Similarity as measured by II gets weaker as the frequency of language-B decreases ( ϵ=0 ). (b) BP theory captures this behaviour by considering fully learned language-A and partially learned language-B, up to level ℓ . (c) The minimum II (across layers) decreases monotonically with the fraction of low-resource language ( PB/P ). (d) Plot adapted from ( Acevedo et al., 2026 ) : Multilingual LLM shows qualitatively similar trend – minimum II between English and other languages decreases monotonically with the online presence of the other language. Both theory and observations display apparent power law behavior in the range probed with exponents α≈0.5 and α≈0.2 respectively.
Figure 5: In both trained transformers and theoretical BP networks, layerwise novelty reveals more strongly similar intermediate computations than raw hidden states.
Figure 6: Comparison of II on raw representations and layer-wise novelty in Mistral-7B-v0.1 and OLMo-3-7B for parallel English–Romanian and English–Polish sentence pairs from Europarl.
Figure 7: Alignment measured by latent cross-probing. Linear probes for latents at level- Lsurface are trained on language-A representations and tested on language-B representations. Accuracy is controlled jointly by the code-switching rate ϵ and the training-set size P : the aligned region (green) appears along the predicted boundary Pthreshold∝ϵ−1 (dashed line).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Schematic of RHM with L=4 . Visible tokens, i.e. leaves, are denoted as xi≡Xi(0) . Level- ℓ parent node of i -th token is denoted as Xi(ℓ) – this introduces a redundancy in notation since this latent can be denoted with any of the 2ℓ subscripts, corresponding to the tokens in its subtree.
Figure 9: Belief propagation for causal next-token prediction. Blue nodes denote upward messages summarizing evidence within observed subtrees; orange nodes denote downward messages along the path to xt+1 . Unfilled nodes are unobserved, and squares denote grammar factors ψℓ . The right panels show the local upward, downward, and top-level updates.
Figure 10: Mean-squared error when predicting layerwise representations from one network using another, for γ=1/2 (a) and γ=0 . (b) BP encodings. Transformer → theory predictions confirm the expected layerwise encoding of BP messages in trained transformers. These heatmaps also match with theory → theory predictions, further solidifying the correspondence. The match is better for γ=1/2 .
Figure 11: Comparing layerwise II profiles for different γ∈[0,1] . Empirical II curves from trained transformers shown for reference.
Figure 12: II for parsimonious implementation of BP where the only latent messages stored correspond to completely observed subtrees – compared with the ordinary BP and trained transformers. The qualitative behaviour is unaffected.
Figure 13: Test loss for transformers trained on MRHM data with ϵ=0.1 and different training set sizes P . The loss scales with P , showing sequential plateaus at the theoretical entropies computed with depth-limited BP.
Figure 14: Transition from encoder-only to encoding-decoding computations in transformers trained on MRHM data with various ϵ and P . Interestingly, the algorithmic transition co-occurs with coordinate alignment (Fig. 7 ).
Figure 15: ϵ=0 version of Fig. 2 in the main text. In this case, the layerwise similarity trends in trained transformers are better captured by that of a theoretical network constructed with an (L+1) -step encoder-only BP algorithm.
Figure 16: Layerwise similarity profiles with varying number of transformer layers. The trend remains qualitatively unchanged, evidencing that transformers stretch the BP computation steps among its layers when more depth is available.
Figure 17: Cross-lingual similarity with varying tree topology. II is shown against relative transformer depth for three splits of a generative hierarchy with L=4 . Across all three splits, II reaches an intermediate-layer minimum and rises toward the output, consistent with the encoder–decoder picture.
Figure 18: Extended version of Fig. 5 . Cross-lingual geometric similarity in trained transformers and theoretical BP networks is measured using II (a) and mKNN (b) for raw representations, layerwise novelty, the residual contribution, and the supra-word residual. Lower II and higher mKNN indicate greater similarity.
Figure 19: Extended version of Fig. 6 . Cross-lingual geometric similarity in Mistral-7B-v0.1, OLMo-3-7B, and EuroLLM-9B is evaluated on parallel English–German, English–Romanian, English–Polish, and English–Finnish sentences from Europarl. II (a) and mKNN (b) are measured for raw representations, layerwise novelty, the residual contribution, and the supra-word residual. Lower II and higher mKNN indicate greater similarity.
As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.
Supantho Rakshit, Adele Goldberg, Henry Conklin
Department of Electrical and Computer Engineering Princeton University · Department of Psychology Princeton University · Laboratory for Artificial Intelligence Princeton University
Large language models exhibit impressive cross-lingual capabilities. However, prior work analyzes this phenomenon through isolated factors and at sparse points during training, limiting our understanding of how cross-lingual generalization emerges--particularly in the early phases of learning. To study the early trajectory of linguistic and translation capabilities, we pretrain a multilingual 1.7B model on nine diverse languages, capturing checkpoints at a much finer granularity. We use word-level translation as a testbed, introducing a novel dataset to trace how translation develops over training through behavioral analyses, model-component analysis, and parameter-based ablations. We find that the model quickly acquires basic linguistic capabilities in parallel with token-level copying, while translation develops in two distinct phases: an initial phase dominated by copying and surface-level similarities, and a second phase in which more generalizing translation mechanisms are developed while copying is refined. Together, these findings provide a fine-grained view of how cross-lingual generalization develops during multilingual pretraining.
Felicia Körner, Maria Matveev, Florian Eichin +3
MaiNLP, Center for Information and Language Processing, LMU Munich, Germany · Munich Center for Machine Learning (MCML) · Department of Mathematics, LMU Munich, Germany +2
We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1) Signal:} We propose that Platonic alignment arises from the universal relationship between objects and attributes, which is encoded linearly in representations according to the Linear Representation Hypothesis (LRH). We provide evidence that LRH helps explain PRH by extracting linear object-attribute features with sparse autoencoders and showing that these sparse representations often exhibit stronger cross-modal alignment than their dense counterparts. {2) Bias:} Models have different implicit biases due to the diverse architectures and training procedures used. We show that this difference can be partially mitigated. Centering and normalization consistently improve cross-model alignment. {3) Noise:} Finite-sample training leads to noise in representations. We provide evidence that representational noise is driven by data scarcity by revealing a strong and consistent positive correlation between word frequency and alignment in LLMs and text embedding models. Synthesizing signal, bias, and noise, we propose a statistical model that refines the Linear Representation Hypothesis and explains further phenomena related to the alignment of representations emerging from diverse modern AI architectures.
Kiril Bangachev, Guy Bresler, Yury Polyanskiy
Department of Electrical Engineering and Computer Science · Massachusetts Institute of Technology