Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.
Figures & tables
Figure 1: Two-dimensional visualization using the first two principal components for the activation ensemble Φr and its decomposed interaction ensembles HABr and HABCr for the prompt template in Section 2.1 using Llama-3.1-8B. Here, γ=B−Amod12 and the expected answer is D=C+γmod12 . Before fitting the principal components, we exclude prompts satisfying A=B , B=C , or C=A (see Appendix B.3 ). The bottom left of each panel shows the coloring according to either γ or D . HABr gets organized by γ in early layers. Until layer 17, HABCr is unstructured but becomes organized by D at layer 18. See Section 3 for a quantitative analysis.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Domain
Prompt template
Llama-3.2-3B
months
Hi {name}, the {noun} is scheduled from {A} {dates} to {B} {dates}. This is the same as time from {C} {dates} to
hours
Hi {name}, the {noun} is scheduled from {day} {A} o’clock to {day} {B} o’clock. This is the same time as from {day} {C} o’clock to {day}
weekdays
Hi {name}, the {noun} is scheduled from {A} to {B}. This is the same as time from {C} to
music
Hi {name}, on the musical scale, the interval from the note {A} to the note {B} is the same as the interval from the note {C} to the note
Llama-3.1-8B
months
Hi {name}, the {noun} is scheduled between {A} {dates} and {B} {dates}. This is the same time as between {C} {dates} and
hours
Hi {name}, the {noun} is scheduled between {day} {A} o’clock and {day} {B} o’clock. This is the same time as between {day} {C} o’clock and {day}
Appendix
Table 1: Prompt templates used for each model and cyclic domain. A,B,C and the replicate variables vary as described in Tables 2 and 3 . N/A indicates domains excluded because the cyclic values do not tokenize to a common width.
Variable
Values
name
Each name should tokenize to one token. There are 194 approved names for each Llama, Qwen, and Mistral model, and 200 for each Gemma model from the list of 200 most common names of the last century ( Social Security Administration, n.d. ) .
Table 3: Values assigned to the cyclic variables A , B , and C in each domain. Months and hours have cycle length N=12 , while weekdays and musical notes have N=7 .
Model
Performance metric
Months
Hours
Weekdays
Music
Llama-3.2-3B
Top-1 accuracy
53.4±2.2 %
78.5±3.2 %
78.3±2.5 %
70.9±0.3 %
Top-3 accuracy
72.2±2.8 %
95.2±1.1 %
96.8±0.8 %
91.9±0.5 %
Llama-3.1-8B
Top-1 accuracy
66.0±3.9 %
79.9±1.8 %
78.8±1.1 %
79.3±0.7 %
Top-3 accuracy
84.8±2.6 %
94.7±1.5 %
98.6±0.3 %
90.2±0.4 %
Qwen-2.5-7B
Top-1 accuracy
62.7±7.0 %
N/A
75.4±2.4 %
89.5±0.6 %
Top-3 accuracy
82.5±6.8 %
N/A
98.7±0.8 %
99.7±0.2 %
Appendix
Table 4: Top-1 and top-3 accuracy by model and domain. Values are mean ± standard deviation across 10 replicates, computed over all prompts, including coincident values of A , B , and C .
Ensemble
0
1
2
3
4
5
6
7
8
9
10
11
HAB
7.445
5.875
3.899
3.568
3.144
2.551
2.391
2.307
2.595
2.700
3.188
4.023
HBC
4.343
1.921
0.918
0.961
0.971
1.021
0.853
0.976
1.022
1.378
1.377
1.575
HCA
6.448
3.013
1.311
1.457
1.478
1.624
1.476
1.487
1.263
1.220
1.589
3.189
Appendix
Table 5: Mean norm of each interaction vector categorized by γ=B−A , γ′=C−B , and γ′′=A−C , respectively, for a single replicate of the months prompt template at layer 15 of Llama-3.1-8B. The interaction vectors associated with zero difference have substantially larger norms than the remaining classes. This motivates the exclusion of coincident-variable prompts from the quantitative analyses in the main text.
Control
Prompt template
Non-cyclic
Hello {name}, the {noun} is scheduled between {A} {dates} and {B} {dates}. This is the same time as between {C} {dates} and
Off by 1
Hello {name}, the {noun} is scheduled between {A} {dates} and {B} {dates}. The month that comes after {C} {dates} is
List
Hello {name}, the talks are on {A} {dates}, on {B} {dates}, and on {C} {dates}. The second talk is on
Appendix
Table 6: Control prompt templates for Llama-3.1-8B.
Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality. In this work, we investigate how instruction-tuned Transformer models (LLaMA and Mistral) encode discourse relations in English, with a particular focus on the contrasting relations of causation and antithesis. Framing the task as a next-token prediction task and applying a suite of interpretability techniques to test model internals, our findings show that certain early layers make predictive decisions at mid-sequence tokens, while some mid-level layers finalize their decisions closer to the last token. Most of the remaining layers primarily propagate earlier decisions rather than actively influencing them. Additionally, we observe that some layers exhibit a preference for one answer over alternatives, suggesting asymmetric representation of discourse-based reasoning.\footnote{Our code is available at https://github.com/abhidipbhattacharyya/causation_vs_antithesis}
Abhidip Bhattacharyya, Shira Wein
University of Massachusetts Amherst · University of South Florida
Because large language models (LLMs) are impressively successful in predicting text, it appears that they must have access to a 'world model' representing causal and definitional structure. However, the dominant formalisms of modern causal inference -- Judea Pearl's interventionist approach and the Neyman-Rubin potential outcomes framework -- struggle to illuminate how LLMs learn causal structure. I resolve this puzzle by arguing that LLMs employ a specific inductive approach based on a difference-making logic -- sometimes called variational induction. I demonstrate how central aspects of this logic are realized during training, where LLMs require enormous amounts of text data from a wide range of contexts to identify difference- and indifference-makers within word sequences. Furthermore, I analyze specific architectural features of LLMs -- such as token embeddings and self-attention -- to determine their roles in variational induction. The difference-making logic of LLMs fundamentally parallels the experimental method, where causal relations are derived by systematically varying individual circumstances to determine their influence on a phenomenon.
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe R2-values from 0.83-0.99 across HMM and LLM combinations, ranging from early to late layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade performance substantially. Together, these results provide representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative model. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale LLMs.
Daniel Balcells, Andrew Jun Lee, Chirag Rastogi +3