Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which layers specialize in this fundamental task or how other tokens in the same sequence, termed context tokens, influence the model's ability to copy the entity tokens. To address these questions, we conduct experiments on Qwen3-8B using two novel methods: genie-in-a-bottle, which controls exactly which layers can participate in an entity-copying task, and attention lobotomy, which cuts off specific tokens' attention to entity tokens without affecting the remaining attention distribution. We find that two distinct groups of layers in the second half of the model are both necessary and sufficient for entity copying. Moreover, in addition to the decoding position's attention to entity tokens, context tokens' attention to entity tokens also proves necessary for copying the exact tokens, even though context tokens do not store entity information themselves unless they satisfy particular semantic properties. Our findings establish the critical role of late layers in entity copying under the guidance of context tokens, calling for future work on how models propagate and consume entity information.
Figures & tables
Figure 1: Two methodologies for studying entity copying. (a) Genie-in-a-bottle confines access to the entity token E to the layer window from ℓstart to ℓfinish by performing three separate forward passes before, within, and after the window. Only the second pass, within the window, enables attention to E . (b) Attention lobotomy removes specific attention pathways x→y to study their effect on entity copying. The effective attention weights for selected destination tokens y are zeroed after softmax.
Figure 2: Two of the five prompt templates, shown with the same entity and expected answer.
Figure 3: The general_success_rate when only a fixed window of attention layers can attend to entity tokens, for different window spans.
Figure 4: Output distribution after removing attention to entities within a fixed window.
Figure 5: Comparison of general_success_rate with attention to the control entity enabled versus disabled outside the designated window.
Figure 6: (a) Breakdown of output categories for each template after removing context tokens’ attention to entities. (b) For the two templates with the most unsuccessful outputs, the decoding position’s attention to entities decreases sharply after context tokens stop attending to them.
Figure 7: (a) Restoring context tokens’ attention to entity tokens in just two layers improves subsequent positions’ copying success. (b) With attention to entity tokens disabled and attention to context tokens intact, the model generates mostly wrong names or non-name responses.
Figure 8: Number of successful outputs under four context conditions with attention to entity tokens removed.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Template
Prompt
DF
Remember this fact: the person’s name is {entity}. Question: What is the person’s name? Respond with only the name. Answer:
PIL
Among apple, mouse, {entity}, and flute, exactly one item is a person’s name. Respond with only that person’s name. Answer:
F
Here is a fact: {entity} is my friend. Question: What is my friend’s name? Respond with only the name. Answer:
VR
The visitor signed the register with the name {entity}. Question: What name did the visitor write? Respond with only the name. Answer:
NB
The name printed on the badge is {entity}. Question: What name is printed on the badge? Respond with only the name. Answer:
Appendix
Table 1: Complete prompt templates. Line breaks within each prompt are preserved.
One token (50 entities)
Two tokens (50 entities)
Einstein
Picasso
Emily Dickinson
Vera Rubin
Newton
Gandhi
Mary Shelley
John Locke
Darwin
Mandela
Oscar Wilde
Adam Smith
Tesla
Lincoln
Victor Hugo
Benjamin Franklin
Euler
Churchill
Toni Morrison
George Washington
Gauss
Thatcher
Barack Obama
Brad Pitt
Appendix
Table 2: Complete entity lists, grouped by entity-span length.
Language models used in retrieval-augmented settings must arbitrate between parametric knowledge stored in their weights and contextual information in the prompt. This work presents a mechanistic study of that choice by extracting an \emph{arbitration vector} from model activations on a curated dataset designed to disentangle (i) irrelevant contexts that elicit parametric recall and (ii) relevant but false contexts that elicit copying. The vector is computed as the residual-stream centroid difference between these regimes across 27 relations, and is injected as an additive intervention at selected layers and token spans to steer behavior in two directions: Copy→Recall (suppressing context use) and Recall→Copy (inducing the model to copy any token from the context). Experiments on three architectures (decoder-only and encoder/decoder) and two open-domain QA benchmarks show consistent behavior shifts under moderate scaling while monitoring accuracy and fluency. Mechanistic analyses of attention routing, MLP contributions, and layer-wise probability trajectories reveal an asymmetry: inducing copying is an easy reactivation'' process that can be triggered at different locations in the input, while restoring recall is a suppression'' process that is more fragile and strongly tied to object-token interventions.
Mehrdad Farahani, Franziska Penzkofer, Richard Johansson
1Chalmers University of Technology · University of Gothenburg
Entity tracking (ET), the ability to keep track of states, is a fundamental skill that underlies complex reasoning. An increasing amount of work investigates how transformer language models (LMs) solve entity binding without state changes. However, there is limited understanding of how non-toy LMs address ET problems of realistic difficulties expressed in natural language. To this end, we investigate the mechanisms underlying ET in more complex scenarios featuring multiple state-changing operations. We find that LMs do not incrementally track world states across tokens or query-relevant states across layers, but simply aggregate relevant information in parallel at the last token when the query becomes evident. We further investigate mechanisms of individual operations (PUT, REMOVE, MOVE) to characterize this non-incremental ET mechanism. Surprisingly, LMs implement the REMOVE operation with a fragile global suppression tag; this global removal mechanism predicts various failure modes that we confirm behaviorally. We provide a mechanistic solution of nullifying this tag to partially address this issue. Overall, our findings reveal that LMs solve a fundamentally sequential task using a non-sequential strategy. More broadly, our work illustrates how behavioral and mechanistic analyses can fruitfully interact. Behavioral results inform mechanistic hypotheses, and insights from mechanistic analyses help build stronger behavioral evaluations by predicting failure modes missing from existing evaluations.
Zilu Tang, Qiao Zhao, Gabriel Franco +4
Department of Computer Science, Boston University, Boston, USA · Department of Data Science, Monash University, Indonesia · Faculty of Computer Science, University of Vienna, Austria +1
Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on long-context data can help but is expensive due to the quadratic scaling of Attention. We observe that most tokens do not require (Global) Attention over the entire sequence and can rely on local context. Based on this, we propose L2A (Learning To Attend), a layer that enables conditional (token-wise) long-range memory access by deciding when to invoke global attention. We evaluate L2A on Qwen 2.5 and Qwen 3 models, extending their effective context length from 32K to 128K tokens. L2A matches the performance of standard long-context training to within 3% while skipping Global Attention for ∼80% of tokens, outperforming prior baselines. We also design custom Triton kernels to efficiently implement this token-wise conditional Attention on GPUs, achieving up to ∼2× improvements in training throughput and time-to-first-token over FlashAttention. Moreover, L2A enables post-training pruning of highly sparse Global Attention layers, reducing KV cache memory by up to 50% with negligible performance loss. Our code is released under Apache 2.0 at https://github.com/awslabs/hybrid-model-factory/tree/main/examples/research/L2A.