Large language models (LLMs) were invented for natural language tasks such as translation, but they have proved that they can perform highly complex functions across domains. Additionally, they have been thought to develop new skills without being trained on them. These learning capabilities lead to LLMs adoption in a wide range of domains. Thus, it is imperative that we understand their operating mechanisms and limitations for proper diagnostics and repair. The earlier studies proposed that high level concepts are encoded as linear directions in LLMs activation space and that the geometry of embeddings have semantic meanings. Inspired by these studies, we hypothesize that LLMs may use subspaces and vector algebra in subspaces to perform tasks. To address this hypothesis, we analyze LLMs' functional modules and residual streams collected from LLMs engaging in in-context learning (ICL), one of the emergent abilities. Our analyses suggest that 1) LLMs can create subspaces, where evidence can be accumulated and 2) ICL tasks can be solved via simple algebraic operations in subspaces.
Figures & tables
Figure 1: ICL prompts. (A), The structure of the used ICL prompts. (B), The structure of a single example ICL with uniform attention score Aij . This panel illustrates how symbols qi , si , ai are mapped onto ICL tokens.
Figure 2: Explained variance accounted for by 30 principal components. x -axis denotes principal components, and y -axis denotes the cumulative explained variance. (A), Explained variance evaluated from LLMs engaging the antonym task. In the panel, 6 models are clarified using different colors (see the inset). (B)-(F), the same as (A), but the task is synonym, country-capital and English-French, country-capital and person-sport, respectively.
Model
ANT
SYN
CC
EF
PC
PS
Neox
(0.97, 0.02)
(0.97, 0.01)
(0.98, 0.01)
(0.97, 0.01)
(0.96, 0.01)
(0.99, 0.02)
Olmo
(0.86, 0.01)
(0.8, 0.01)
(0.89, 0.01)
(0.9, 0.01)
(0.85, 0.02)
(0.95, 0.02)
Pyt
(0.96, 0.01)
(0.97, 0.01)
(0.99, 0.02)
(0.97, 0.01)
(0.98, 0.01)
(0.99, 0.01)
GPT
(0.97, 0.01)
(0.95, 0.01)
(0.98, 0.01)
(0.96, 0.01)
(0.96, 0.01)
(0.97, 0.01)
Llma
(0.94, 0.01)
(0.81, 0.01)
(0.96, 0.01)
(0.88, 0.01)
(0.85, 0.01)
(0.97, 0.02)
Gemma
(0.91, 0.01)
(0.91, 0.01)
(0.92, 0.01)
(0.65, 0.01)
(0.89, 0.01)
(0.94, 0.02)
Table 1: Comparison of R2 with the baselines. We compares the mean values ( V1 ) of top-3 from our original experiments and those ( V2 ) from the control experiments. For each model and task, we display ( V1 , V2 ). That is, the first number in the parenthesis denotes the mean value of top 3 R2 from the original experiment, whereas the second number denotes the mean value of top 3 R2 from the control experiment. ANT, SYN, CC, EF, PC, PS stand for antonym, synonym, country-capital, English-French, Product-Campany and Person-Sport, respectively. Neox, Olmo, Py, GPT, Llma and Gamma denote GPT-neox-20B, OLMo-2- 0325 -32B, pythia-12B Llma-3.1-8B and gemma-3-27b-it, respectively.
Figure 3: Quality of linear regression evaluated. R2 is evaluated using LLMs engaging in ‘antonym’ task. x -axis denotes the transformer layer, and y -axis denotes the principal components. (A)-(F) shows R2 estimated from GPT-j-6B, Meta-Llama-3.1-8B, OLMo-2- 0325 -32B, Phythia-12B, gemma-3-27b-it and GPT-NEOX- 20 B.
Figure 4: The same as Fig. 3 , but the task is country-capital.
Figure 5: Query (que), separator (sep) and answer (ans) tokens in subspace spanned by 3 principal components with the highest R2 s. Red, green and blue dots represent separator, answer and query tokens, respectively (see the inset). All tokens are collected from 6 LLMs engaging in the antonym task. The model is specified by the name above each plot.
Model
PS
CC
PC
pythia-12B
(1.0, 0.1)
(1.0, 0.25)
(1.0, 0.16)
gpt-neox-20b
(1.0, 0.09)
(0.96, 0.08)
(1.0, 0.18)
gpt-j-6B
(0.98, 0.12)
(0.99, 0.01)
(1.0, 0.08)
gemma-3-27b-it
(0.7, 0.02)
(0.72, 0.0)
(0.09, 0.06)
Table 2: This table shows how often LLMs change their answers when the last token is perturbed via top-3 and bottom-3 components. For each case, the first number in the parenthesis shows the ratio of the trials, in which LLMs change their answers when the perturbation is delivered via top-3 components. The second number shows the ratio of the trials, in which LLMs change their answers when perturbation is delivered via bottom-3 components. PS, CC and PC stand for Person-Sport, Country-Capital and Product-Company, respectively.
Figure 6: The same as Fig. 5 , but the task is country-capital.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Supplementary Figure 1: Comparison of query ql in Panel (A), separator sl in Panel (B), answer al in Panel (C), and Δ in Panel (D) when LLMs make the correct predictions and when they make incorrect ones. In this experiment, we use t -test and report p -values. The blue, orange and green lines denote country-capital task, product-company task and person-sport task, respectively.
Supplementary Figure 2: The same as Fig. 3 , but the task is the English-French.
Supplementary Figure 3: The same as Fig. 3 , but the task is synonym.
Supplementary Figure 4: The same as Fig. 3 , but the task is Product-Company.
Supplementary Figure 5: The same as Fig. 3 , but the task is Person-Sport.
Supplementary Figure 6: The same as Fig. 5 , but the task is English-French.
Supplementary Figure 7: The same as Fig. 5 , but the task is synonym.
Supplementary Figure 8: The same as Fig. Fig. 5 , but the task is product-company.
Supplementary Figure 9: The same as Fig. Fig. 5 , but the task is person-sport.
Model
ANT
SYN
CC
EF
PC
PS
Neox
(0.97, 0.04)
(0.97, 0.04)
(0.98, 0.05)
(0.97, 0.04)
(0.96, 0.08)
(0.99, 0.06)
Olmo
(0.86, 0.05)
(0.8, 0.03)
(0.89, 0.02)
(0.9, 0.08)
(0.85, 0.02)
(0.95, 0.02)
Pyt
(0.96, 0.03)
(0.97, 0.04)
(0.99, 0.03)
(0.97, 0.02)
(0.98, 0.09)
(0.99, 0.02)
GPT
(0.97, 0.06)
(0.95, 0.04)
(0.98, 0.03)
(0.96, 0.03)
(0.96, 0.09)
(0.97, 0.05)
Llama
(0.94, 0.06)
(0.81, 0.04)
(0.96, 0.02)
(0.88, 0.03)
(0.85, 0.05)
(0.97, 0.03)
Genna
(0.91, 0.03)
(0.91, 0.03)
(0.92, 0.02)
(0.65, 0.02)
(0.89, 0.04)
(0.94, 0.02)
Appendix
Supplementary Table 1: We compare the mean values ( V1 ) of top-3 R2 from our original experiments and those ( V2 ) from the control experiments. For each model and task, we display ( V1 , V2 ). That is, the first number in the parenthesis denotes the mean value of top-3 from the original experiments, whereas the second number denotes those from the control experiments. ANT, SYN, CC, EF, PC, PS stand for antonym, synonym, country-capital, English-French, Product-Campany and Person-Sport, respectively.Neox, Olmo, Pyt, GPT, Llama and Gamma denote GPT-neox-20B, OLMo-2- 0325 -32B, pythia-12B Llma-3.1-8B and gemma-3-27b-it, respectively.
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe R2-values from 0.83-0.99 across HMM and LLM combinations, ranging from early to late layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade performance substantially. Together, these results provide representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative model. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale LLMs.
Daniel Balcells, Andrew Jun Lee, Chirag Rastogi +3
Large language models (LLMs) exhibit remarkable flexibility: they can adapt to novel tasks from in-context examples without any parameter updates, a capability known as in-context learning (ICL). Prior work on synthetic tasks has shown that ICL can implement specific algorithms, demonstrating architectural competence, and mechanistic analyses have identified key circuits that support this behavior. However, because in-context computation -- regardless of its algorithmic form -- relies on transformations in high-dimensional representation space, it remains unclear how the geometry of that space shapes ICL effectiveness. Motivated by the neuroscience view of classification as the untangling of neural representations, we hypothesize that ICL depends on the successful online untangling of task-relevant representations. To test this idea, we study how LLMs classify in-context examples whose labels are defined by the model's own internal representations with known structure. We show that ICL performance correlates systematically with the representational structure of the underlying classification task and that successful ICL is accompanied by geometric reorganization that increases online separability. We further find that LLM behavior is well described by a prototype-like algorithm that integrates evidence while reshaping representations to support classification. These findings offer a geometric account of ICL in pretrained LLMs, establish representational geometry as a mechanistic constraint on ICL, and quantify the gap between what pretrained representations afford and what in-context learning can exploit.
Hua-Dong Xiong, Li Ji-An, Robert C. Wilson +2
School of Psychological and Brain Sciences, Georgia Tech · Department of Psychology, New York University · School of Psychological and Brain Sciences, Georgia Tech 3 Center of Excellence for Computational Cognition, Georgia Tech +2
Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along single linear directions, but recent findings suggest complex behaviours, such as the refusal to answer harmful queries, live in multi-dimensional subspaces. However, existing methods for extracting these subspaces are computationally expensive, which becomes prohibitive on reasoning models who produce long reasoning traces. By adapting the Recursive Feature Machine (RFM) algorithm -- which can be computed efficiently -- with a probe-informed initialization, we are able to identify the multi-dimensional refusal subspace in seconds, on reasoning (Qwen 3) and non-reasoning (Qwen 2.5) models. While RFM allows for faster subspace identification, it also showed better performances on the ablation task than its alternatives. More work is planned to better understand the relations between subspaces found by different methods. If confirmed, RFM could be a cheap and scalable complement to existing subspace-extraction methods in LLMs.
Thomas Winninger
Télécom SudParis, Évry-Courcouronnes, France · ENS Paris-Saclay, Gif-sur-Yvette, France