Signatures of semantic search in the activations of large language models
Authors: Luke Leckie, Peter M. Todd, Jacob G. Foster
Organizations: Department of Informatics, Luddy School of Informatics, Computing, and Engineering, Indiana University, Bloomington, IN · Cognitive Science Program, Indiana University, Bloomington, IN · Center for Possible Minds, Indiana University, Bloomington, IN · Department of Psychological and Brain Sciences, Indiana University, Bloomington, IN · Santa Fe Institute, Santa Fe, NM
When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and between-cluster switching ("explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., "water") increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.
Figures & tables
Figure 1: Semantic fluency and LLM J-lens activations. A. Left: Example Semantic Fluency Task (SFT) output. Cluster labels are given by colours and switches between clusters are indicated by vertical black lines in-between distinct clusters. Right: Uniform Manifold Approximation (UMAP) of Word2Vec embeddings of animal names. Colours indicate extended Troyer category designations. B. Illustration of LLM J-lens activations during SFT. Left neural network represents an LLM. Beige-highlighted layer represents residual stream representations of a layer under study by the J-lens. Beige circle containing concepts represents a set of concepts included within the J-space for that layer. Note that in our experiments we considered the J-space as the top-25 strongest J-lens activations. J-lens activations are indicated by numbers to the right of concepts. Items that have already been verbalised are indicated by strikethroughs. Higher-level abstract category-related labels (“water” and “pet”) are indicated in blue and red font, respectively. The in-category J-space supply is computed as the number of items (not including category-related labels) within the J-space that belong to the category from which the model is presently naming items but that have not been verbalised yet. In this example, as the LLM names water animals and depletes the in-category J-space supply for that category, it switches to naming animals within the pet category. Note that the J-lens activation for the abstract category-related “pet” label increases in anticipation of switching to that category (See Figure 3 for analysis).
Figure 2: Semantic fluency item activations and J-space depletion. A. Comparison of next-token probabilities for items participating in clustering versus switching events across models. Coloured bars extend to the average difference in next token probabilities for switching events minus clustering events. Black whiskers extend to ± standard error (se). Next-token probabilities were significantly lower during switching (linear models, all p < 0.0001***) for all models. B. Z-scored J-lens activations for items participating in switching, measured in advance (negative lag) and after (positive lag) their emission (lag=0). Points represent means across seeds and layers; shaded areas represent standard errors. C. Δ in Z-scored J-lens activations across fractionalised layers, as measured immediately prior to item emission (lag=-1). Points represent means across seeds and bars around points extend to ± se. Stars for means indicate significantly different activations (p < 0.05 after Benjamini-Hochberg [BH] correction) of item activations between items participating in switching versus clustering, as tested by linear models. D. The probability of switching to a new category as a function of the in-category J-space supply, averaged across layer blocks and models. More yellow points indicate earlier layer blocks and purple, later. Note that J-space supply decreases from left to right.
Figure 3: J-lens activations of category-related labels during LLM semantic fluency. A. Δ in the J-lens activations of category labels, subtracting the mean activation of each label during production of items in unrelated categories from mean activation during production of items in clusters with that category label (target category). Points indicate means across all tested labels and seeds. Bars extend to ± se. B-D. Layer-averaged J-lens activations of category labels preceding (negative x-values) and following a switch (positive x-values) into the target category ( B ) or an unrelated category ( C ), as well as the Δ between the two (target minus unrelated activation) ( D ). Points indicate means across seeds and layers; shaded areas indicate standard errors. A and D : Stars for means indicate significantly different activations (p < 0.05 after BH correction) of category-related labels between target and unrelated categories. E. Δ in target minus unrelated category label activations for each tested label (x-axis) for Gemma-2-9B, averaged across layers and seeds. More positive values (lighter cells) indicate stronger Δ activation of that label. Note that the first instance of “domestic” maps to the “domestic animals” category and the second maps to the “pets” category.
Figure 4: Steering model activations to induce switching into target categories. A. Models were induced to switch into target categories by steering vectors generated from averaging activations over item-level representations or from individual abstracted category-related labels. B-C. Effect of steering using item ( B ) or category-related ( C ) steering vectors. Points represent means across target-level probabilities computed across seeds and target categories tested and bars extend to ± se.
Figure 5: Identifying and steering a generic direction for switching in model activation space. A. Position of anticipatory and event tokens during SFT emission. Bottom neural net: anticipatory steering vectors, weighted by s , are used to bias switching. Positive s should bias increased switching and negative s , decreased. B-C. AUC from logistic regression models applied across layers and models for predicting switching versus clustering based on anticipatory ( B ) and event activations ( C ). D. Layer-averaged loading of model activations onto the anticipatory residual stream direction during clustering, up until switching. Anticipatory activations progressively ramp nearing a switch. Note that Qwen (orange) has the smallest mean cluster size in our experiments (2.4 ± 2.8 ). E-F. Average effect of positive ( E ) and negative steering ( F ) on switching across models and layers, relative to the average switch rate for each model. Points represent means across 100 SFTs. Stars for means indicate significant effects of steering as compared to noise. G. SFT output from Gemma-2-9B under baseline and each perturbation at layer 26.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: A. Z-scored J-lens activations for items participating in clustering, measured in advance (negative lag) and after (positive lag) their emission (lag=0). Points represent means across seeds and layers and shaded areas, standard errors. C. Δ in Z-scored J-lens activations for switching minus clustering. Points represent means and shaded areas, standard errors (note that standard errors are very narrow). Stars indicate significant comparisons between switch- and cluster- associated activations.
model
layer
beta
F
df
p
sig
Gemma-9B
2
-0.187
242.080
1
2.84E-28
TRUE
Gemma-9B
4
-0.244
263.261
1
1.51E-29
TRUE
Gemma-9B
6
-0.176
130.854
1
8.64E-20
TRUE
Gemma-9B
8
-0.259
544.888
1
8.83E-42
TRUE
Gemma-9B
10
-0.415
951.408
1
4.04E-52
TRUE
Gemma-9B
12
-0.411
1011.960
1
2.94E-53
TRUE
Appendix
Table S1: Statistics comparing J-lens activations of items participating in clustering or switching across layers (Fig. 2C).
model
item lag
beta
F
df
p
sig
Gemma-9B
-10
0.009
0.245
1
6.51E-01
FALSE
Gemma-9B
-9
0.016
0.768
1
4.47E-01
FALSE
Gemma-9B
-8
-0.066
11.302
1
1.93E-03
TRUE
Gemma-9B
-7
-0.079
16.463
1
2.08E-04
TRUE
Gemma-9B
-6
-0.148
59.068
1
4.72E-11
TRUE
Gemma-9B
-5
-0.124
42.628
1
8.58E-09
TRUE
Appendix
Table S2: Statistics comparing J-lens activations of items participating in clustering or switching across lag values from their emission, averaged over layers (Fig. S1B).
model
depth range
r
t
df
p
sig
Gemma-9B
0.00-0.20
-0.367
-30.358
1
1.47E-51
TRUE
Gemma-9B
0.20-0.40
-0.402
-26.220
1
2.99E-46
TRUE
Gemma-9B
0.40-0.60
-0.390
-31.876
1
3.66E-53
TRUE
Gemma-9B
0.60-0.80
-0.362
-17.728
1
1.74E-32
TRUE
Gemma-9B
0.80-1.00
-0.306
-27.619
1
4.29E-48
TRUE
Gemma-2B
0.00-0.20
-0.336
-18.451
1
1.19E-33
TRUE
Appendix
Table S3: Statistics to accompany Fig. 2D. Pooled correlation coefficients (r) are tested against 0 with one-sample t-tests.
model
depth range
beta
F
df
p
sig
Gemma-9B
0.00-0.20
-0.091
82.890
1
1.64E-14
TRUE
Gemma-9B
0.20-0.40
-0.137
131.235
1
1.89E-19
TRUE
Gemma-9B
0.40-0.60
-0.140
202.842
1
5.26E-25
TRUE
Gemma-9B
0.60-0.80
-0.059
24.841
1
3.31E-06
TRUE
Gemma-9B
0.80-1.00
-0.016
1.987
1
1.62E-01
FALSE
Gemma-2B
0.00-0.20
-0.039
3.717
1
9.46E-02
FALSE
Appendix
Table S4: Statistical comparisons testing the effect of J-space in-category depletion on switching probability, as compared to depletion of the overall category pool.
model
J-space size
depth range
beta
F
df
p
sig
Gemma-9B
5
0.00-0.20
0.033
4.530
1
1.00E+00
FALSE
Gemma-9B
5
0.20-0.40
-0.061
15.266
1
2.85E-04
TRUE
Gemma-9B
5
0.40-0.60
-0.060
16.248
1
2.73E-04
TRUE
Gemma-9B
5
0.60-0.80
-0.069
17.788
1
2.73E-04
TRUE
Gemma-9B
5
0.80-1.00
0.064
16.907
1
1.00E+00
FALSE
Gemma-9B
15
0.00-0.20
-0.043
16.140
1
1.43E-04
TRUE
Appendix
Table S5: Robustness analyses testing the effect of J-space size on the comparisons reported in S4 for Gemma models. Sizes of 5, 15, and 25 were tested.
model
item
beta
F
df
p
sig
Gemma-9B
-6
0.107
23.676
1
4.31E-06
TRUE
Gemma-9B
-5
0.123
61.405
1
6.35E-12
TRUE
Gemma-9B
-4
0.154
71.636
1
3.13E-13
TRUE
Gemma-9B
-3
0.113
41.314
1
4.97E-09
TRUE
Gemma-9B
-2
0.266
165.96
1
1.12E-22
TRUE
Gemma-9B
-1
0.158
344.957
1
1.08E-33
TRUE
Appendix
Table S6: Statistics for Figure 3D, comparing the activation of category-related labels before and after switching into a target category, versus an unrelated category
model
beta
chi2
p
sig
Gemma-9B
-0.316
525.373
1.43E-115
TRUE
Gemma-2B
-0.239
179.959
1.24E-40
TRUE
Qwen-7B
-0.105
30.040
5.29E-08
TRUE
Llama-3B
-0.079
20.439
6.16E-06
TRUE
Llama-8B
-0.104
67.895
2.87E-16
TRUE
Appendix
Table S7: Statistics to accompany Figure 5D. Linear models testing the association between number of items until a cluster and activation of anticipatory loading
model
layer
base switch rate
steer switch rate
noise switch rate
p
Gemma-9B
2
0.29
0.27
0.27
1.00E+00
Gemma-9B
4
0.29
0.32
0.26
1.15E-12
Gemma-9B
6
0.29
0.32
0.27
1.93E-11
Gemma-9B
8
0.29
0.32
0.28
4.34E-09
Gemma-9B
10
0.29
0.36
0.30
5.06E-11
Gemma-9B
12
0.29
0.45
0.34
1.15E-31
Appendix
Table S8: Statistics to accompany Figure 5E. Fisher tests comparing switch rates during positive steering against noise.
model
layer
base switch rate
steer switch rate
noise switch rate
p
Gemma-9B
2
0.29
0.26
0.28
3.30E-02
Gemma-9B
4
0.29
0.25
0.28
1.51E-05
Gemma-9B
6
0.29
0.25
0.29
1.70E-06
Gemma-9B
8
0.29
0.26
0.28
1.15E-01
Gemma-9B
10
0.29
0.28
0.32
4.17E-05
Gemma-9B
12
0.29
0.32
0.32
7.75E-01
Appendix
Table S9: Statistics to accompany Figure 5F. Fisher tests comparing switch rates during negative steering against noise
Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By applying trajectory-based NLP metrics to the items generated by 82 human participants and LLM output across eight temperature settings, we quantified three complementary dimensions: entropy (step size predictability), distance to next (successive semantic steps), and distance to centroid (global dispersion). Humans exhibited higher entropy, larger semantic steps and broader dispersion than all LLMs, indicating more variable and exploratory search. Temperature tuning produced only partial alignments, as individual metrics matched between humans and LLMs at specific settings, but no configuration reproduced the complete human profile (in all dimensions). These findings suggest that human semantic search implements a distinctive balance between local exploitation and global exploration that current model architectures fail to reproduce.
Gabriel Paris-Colombo, Rodrigo M. Cabral-Carvalho, Felipe D. Toro-Hernández
Center for Mathematics, Computing and Cognition, Federal University of ABC
Recent work shows that large language models (LLMs) encode behavioural traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioural steering. Building on this, we treat them as dynamic signals instead: probes we can monitor and intervene on as reasoning unfolds. We use the term polylogue to denote the time series of alignments between persona vectors and hidden activations over the course of generation. Experiments across four open-weight models show that polylogue features predict correctness on MMLU-Pro competitively with low-dimensional activation baselines, while remaining interpretable through their associated persona directions. They also suggest concrete steering targets, namely which latent directions to modulate at different stages of a response. We instantiate this as a simple paragraph-conditioned intervention that improves accuracy on three of four models, pointing to stage-aware latent steering as a promising direction for reasoning-time control. Together, this positions the polylogue as an interpretable tool for reasoning-time monitoring and intervention.
Nils A. Herrmann, Leander Girrbach, Kirill Bykov +1
1Technical University of Munich · 2Helmholtz Munich · 3Munich
Large language models (LLMs) encode rich concept-like information, but represent it implicitly through distributed statistical associations rather than as explicit, structured, compositional concepts. Consequently, concept-level structure is typically \emph{found} rather than \emph{designed}: it is recovered after training through probing or dictionary learning, with no architectural guarantee of stability, compositionality, controllability, or alignment with human conceptual organization. We organize concept-aware interventions along two dimensions: whether concept structure is internally induced or externally grounded, and the stage of the pipeline where it is introduced. This taxonomy reveals three broad patterns: inference-time approaches remain comparatively underexplored, related ideas have developed largely in isolation across pipeline stages, and externally grounded methods span the entire pipeline despite often being described under different terminology. Together, these observations motivate moving beyond recovering concept-like structure from trained models toward designing LLMs with explicit conceptual representations.