cs.LGSep 27, 2026

Rethinking Contextualization by Reinterpreting Attention Head Channels

Authors: Hakaze Cho, Haolin Yang, Zhun Sun, Naoya Inoue, Benjamin Heinzerling, Kentaro Inui

Organizations: RIKEN · Tohoku University · New York University · JAIST · MBZUAI

Abstract

Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete dictionary, lacking a global view of their general behavior. Therefore, we propose a general principle: Globally, we find and estimate that different words carry different amounts of information, and less-informative words tend to absorb more contextual information. Specifically, these low-information words do not absorb contextual words uniformly, and finer-grained selectivity enables more precise routing to promote information transmission between matched words. Moreover, to find what mechanism causes such processing, we reinterpret attention heads as channels gated by their singular vectors and find that: (1) these singular vectors point to the hidden states of more informative words, allowing such words to write their information to others more strongly to act as information sources, and vice versa; and (2) these singular vectors can be viewed equally as hidden state features, enabling automated interpretation of attention heads beyond prior heuristic head discovery, also embedding heads into a continuous space rather than treating them as discrete, independent dictionary entries.

Figures & tables

Appendix figures & tables47 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation

    Sep 23, 2026Paweł Mąka, Yusuf Can Semerci, Jan Scholtes +1Token-Level EntropyDisambiguation

  2. Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation

    Sep 14, 2026Micah Adler, John W. Byers, Mark CrovellaLanguage ModelingMean-Field Limit

  3. Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

    Sep 8, 2026Sara Rizwan, Samaanah Abdus SalamEfficient Long-Context Inference