cs.CLSep 28, 2026

Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context

Authors: Muyu He, Yuchen Liu, Ran Tao, Li Zhang

Organizations: Independent · University of Pennsylvania · Drexel University

Abstract

Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which layers specialize in this fundamental task or how other tokens in the same sequence, termed context tokens, influence the model's ability to copy the entity tokens. To address these questions, we conduct experiments on Qwen3-8B using two novel methods: genie-in-a-bottle, which controls exactly which layers can participate in an entity-copying task, and attention lobotomy, which cuts off specific tokens' attention to entity tokens without affecting the remaining attention distribution. We find that two distinct groups of layers in the second half of the model are both necessary and sufficient for entity copying. Moreover, in addition to the decoding position's attention to entity tokens, context tokens' attention to entity tokens also proves necessary for copying the exact tokens, even though context tokens do not store entity information themselves unless they satisfy particular semantic properties. Our findings establish the critical role of late layers in entity copying under the guidance of context tokens, calling for future work on how models propagate and consume entity information.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. To Copy or Not to Copy: Copying Is Easier to Induce Than Recall

    Jan 17, 2026Mehrdad Farahani, Franziska Penzkofer, Richard JohanssonFactual RecallPrecision Recall

  2. Do Language Models Track Entities Across State Changes?

    May 28, 2026Zilu Tang, Qiao Zhao, Gabriel Franco +4EntitiesState-Tracking

  3. Learning When to Attend: Conditional Memory Access for Long-Context LLMs

    Mar 18, 2026Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4Efficient Long-Context InferenceTime-To-First-Token