How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.
Figures & tables
Figure 1: Useful-token hypothesis. Each attention operation has an unknown useful set whose size K(C) depends on context. Attention and contribution scores give rankings from which sets of different sizes are selected. Ranking inversions occur when tokens outside the useful set score at least as highly as useful tokens, so recovering the useful set may require retaining more than K(C) tokens.
Figure 2: Attention-selected sets exhibit stronger Euclidean separation than random sets across all four models. Geometric precision (left), recall (middle), and extremal FN (right) on OpenWebText. Solid curves show selection by attention weight, and dash-dotted curves show random sets of the same size, each evaluated relative to its own aggregate. Measurements use the final query of 1024-token contexts and are averaged over heads, layers, and documents.
Figure 3: Attention and contribution selection produce lower NLL degradation than random selection at the same attention set sizes. Relative increase in mean NLL over the full model on OpenWebText with 1k-token contexts. Panels compare selection by attention weight (left), contribution magnitude (middle), and random selection (right). For a chosen loss tolerance, the crossing of a curve with that degradation level gives an approximate effective attention set size N∗ .
Figure 4: Stronger geometric separation often accompanies greater NLL degradation. Geometric FN and relative NLL degradation under attention-weight selection, using (a) Euclidean and (b) cosine geometry. Each point corresponds to one set size N , indicated by color, with both quantities averaged across OpenWebText and WikiText-103. Lines connect successive set sizes, and ρ denotes the Spearman correlation along each trajectory.
Figure 5: Longer context improves full-model prediction and increases the effective attention set size. (a) Estimated set size at a 5% relative NLL tolerance against the full model at each context length. (b) Mean full-model NLL. Both panels use 50 OpenWebText documents per model, scoring the same final 128 target tokens at every context length. Shading in (a) shows 95% paired-document bootstrap intervals conditional on the evaluated set sizes.
Figure 6: Longer background text tends to lower full-model accuracy and increase the effective attention set size despite fixed annotated support. BABILong qa1 results on 100 matched examples, with six candidate answers scored using first-continuation-token logits. (a) Full-model candidate accuracy. Shading shows 95% Wilson intervals, and the dotted line indicates chance accuracy. (b) Estimated effective attention set size under attention-weight selection at a 0.10 -nat mean answer-loss tolerance relative to the full model at each background. Open triangles indicate that no evaluated set size through 256 meets the tolerance.
Figure 7: Additional background pushes support tokens down the attention ranking and reduces their attention mass and recall. Measurements at the answer query in BABILong qa1 show support displacement, attention mass, and recall among the 64 highest-weight tokens. Statistics use 86 examples per background after tokenizer-span validation.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Attention-selected sets show stronger separation than random sets on OpenWebText. Geometric precision, recall, and extremal FN under Euclidean and cosine distance. Solid curves show attention ranking and dash-dotted curves show random sets of the same size, each evaluated around its own aggregate. Measurements use the final query of 1024-position windows and are averaged over heads, layers, and documents. The Euclidean advantage is larger.
Figure 9: Contribution ranking reproduces the geometric pattern on OpenWebText. Precision, recall, and extremal FN under Euclidean and cosine distance. Solid curves show contribution ranking and dash-dotted curves show random selection. The evaluation and aggregation follow Figure 8 .
Figure 10: Attention-selected geometry is consistent across corpora. WikiText-103 precision, recall, and extremal FN under Euclidean and cosine distance. Solid curves show attention ranking and dash-dotted curves show random selection. The evaluation and aggregation follow Figure 8 .
Figure 11: Contribution-selected sets also exhibit geometric separation on WikiText-103. Precision, recall, and extremal FN under Euclidean and cosine distance. Solid curves show contribution ranking and dash-dotted curves show random selection. The evaluation and aggregation follow Figure 8 .
Figure 12: Contribution ranking gives similar geometry–loss trajectories. Each point corresponds to a selected-set size N , with geometric FN and relative NLL degradation averaged across OpenWebText and WikiText-103. Color denotes N , lines connect successive sizes, and ρ is the Spearman correlation along each trajectory. Geometry is measured at the final query, while loss uses interventions throughout the model.
Model
Geometry
N≥1
N≥2
N≥8
Qwen-2.5-7B
Euclidean
0.817
0.738
0.371
Qwen-2.5-7B
Cosine
0.983
0.976
0.943
Gemma-7B
Euclidean
0.783
0.690
0.257
Gemma-7B
Cosine
0.433
0.190
-0.943
Llama-3-8B
Euclidean
0.700
0.571
-0.029
Llama-3-8B
Cosine
0.550
0.357
-0.543
Appendix
Table 1: Attention-ranking Spearman correlations between geometric FN and relative NLL degradation after excluding small selected sets. The columns use 9, 8, and 6 measured powers-of-two sizes, respectively.
OpenWebText
WikiText-103
Model
Ranking
N1%∗
N5%∗
N10%∗
N1%∗
N5%∗
N10%∗
Qwen-2.5-1.5B
attention weight
99
39
24
108
42
26
Qwen-2.5-1.5B
contribution magnitude
104
41
25
112
43
27
Qwen-2.5-7B
attention weight
61
33
21
63
31
21
Qwen-2.5-7B
contribution magnitude
59
31
21
63
30
20
Gemma-7B
attention weight
283
226
206
253
195
173
Appendix
Table 2: Effective attention set-size estimates by corpus for 1024-position windows. Deletion is applied at every query, head, and layer. Entries are locally refined crossings at the stated relative NLL tolerance. Gemma-2-9B was evaluated only at 5% and 10%, so its 1% entries are unmeasured.
Figure 13: Matched-target set-size estimates increase on both corpora. Effective attention set size at a 5% relative NLL tolerance, using 50 documents per model and corpus. The same final 128 target tokens are scored at every length within each model–corpus pair. Bands are 95% paired-document bootstrap intervals conditional on the evaluated sizes. Both axes use logarithmic spacing.
Model
Corpus
256
512
1024
2048
Qwen-2.5-7B
OWT
16
31
43
78
Qwen-2.5-7B
WikiText
17
31
45
70
Gemma-7B
OWT
76
150
280
504
Gemma-7B
WikiText
71
137
244
423
Llama-3-8B
OWT
10
17
32
57
Llama-3-8B
WikiText
10
18
34
66
Appendix
Table 3: Attention-ranked effective set-size estimates at a 5% relative NLL tolerance under matched final-128-target scoring. Column headings give context length L . Only complete model–corpus pairs are included. OWT denotes OpenWebText and WikiText denotes WikiText-103.
Figure 14: The retained fraction decreases over the measured range. Attention-ranked N∗/L at a 5% relative NLL tolerance under matched final-target scoring. Bands have the same conditional bootstrap interpretation as in Figure 13 . A decreasing fraction is also compatible with affine growth having a positive intercept.
Model
Corpus
Rank
256
512
1024
2048
Qwen-2.5-7B
OWT
Attn.
27/16/12
46/31/19
73/43/29
164/78/50
Qwen-2.5-7B
OWT
Contr.
26/16/11
47/28/19
71/42/29
158/74/49
Qwen-2.5-7B
WT
Attn.
28/17/13
59/31/22
88/45/31
137/70/49
Qwen-2.5-7B
WT
Contr.
27/16/12
57/31/21
84/45/30
141/69/49
Gemma-7B
OWT
Attn.
106/76/68
192/150/136
340/280/256
646/504/463
Gemma-7B
OWT
Contr.
105/76/68
190/150/136
346/279/256
635/503/461
Appendix
Table 4: Matched final-128-target set-size estimates for the three complete models. Each cell gives N1%∗/N5%∗/N10%∗ , conditional on evaluated sizes. Column headings give L . OWT denotes OpenWebText and WT denotes WikiText-103. Attn. and Contr. denote attention and contribution ranking. The sampling protocol differs from Table 2 .
Model
Corpus
Rank
256
512
1024
2048
Qwen-2.5-7B
OWT
Attn.
22/13/9
35/20/14
55/31/21
102/51/34
Qwen-2.5-7B
OWT
Contr.
21/12/9
34/19/13
54/30/20
100/49/33
Qwen-2.5-7B
WT
Attn.
23/13/9
39/20/14
61/31/22
94/48/34
Qwen-2.5-7B
WT
Contr.
23/12/9
38/20/14
61/30/21
93/48/34
Gemma-7B
OWT
Attn.
89/67/59
157/125/111
275/227/204
483/394/358
Gemma-7B
OWT
Contr.
88/67/59
157/125/111
276/227/203
484/395/358
Appendix
Table 5: All-token set-size estimates from the matched-suffix evaluation. Each cell gives N1%∗/N5%∗/N10%∗ , conditional on evaluated sizes. Column headings give L . Abbreviations follow Table 4 . All predicted positions in each window are scored, so the target population changes with L .
Figure 15: Effective set size depends on which targets are scored. Matched final-target and all-token scoring use the same forward evaluations but average different target populations. The final-target protocol fixes the last 128 targets across lengths within each model–corpus pair.
Figure 16: Additional natural context improves full-model NLL. NLL on the same final 128 targets at each length. The Mistral/OpenWebText curve includes full-model measurements from its partial L=2048 evaluation. Those measurements do not imply a completed set-size estimate.
Figure 17: Mistral-Small-24B retains low NLL degradation with restricted attention. Adaptive refinement of relative NLL degradation for the 8-bit-weight checkpoint. Solid and dashed curves show OpenWebText and WikiText-103. Horizontal lines indicate 1%, 5%, and 10% tolerances.
Figure 18: Pairwise support ranking improves while total displacement grows. Mean empirical probability that a non-support context token scores at least as highly as a support token at the answer query. Background growth increases the number of competitors while this probability decreases. The 0K condition already includes non-support tokens.
Model
0K
1K
2K
4K
Qwen-2.5-7B
32
16
32
64
Llama-3-8B
32
128
256
316
Gemma-7B
64
256
>256
760
Mistral-7B
8
64
256
>256
Appendix
Table 6: Attention-ranked BABILong qa1 set-size estimates at a 0.10-nat mean candidate-answer-loss tolerance relative to each background’s full model. Adaptive extensions are included. An entry >256 denotes no passing evaluated size through 256, with untested integers unresolved.
Model
0K
1K
2K
4K
Qwen-2.5-7B
0.690
1.178
1.303
1.385
Gemma-7B
0.647
0.882
0.933
1.150
Llama-3-8B
0.740
0.773
0.857
1.163
Mistral-7B-v0.3
0.792
1.018
1.151
1.311
Appendix
Table 7: Full-model qa1 candidate-normalized answer loss in nats on 100 matched examples per background. These baselines are used for the corresponding answer-loss degradation measurements.
Figure 19: Full-model answer loss increases with background. Candidate accuracy and mean candidate-normalized answer loss before intervention. Accuracy bands are 95% Wilson intervals, and loss bands use bootstrap resampling over examples. Dotted references indicate uniform six-candidate performance. Accuracy can vary nonmonotonically despite increasing loss.
Model
qa1
qa2
qa3
Qwen-2.5-7B
0.70
0.49
0.28
Gemma-7B
0.71
0.49
0.30
Llama-3-8B
0.67
0.49
0.26
Mistral-7B-v0.3
0.63
0.52
0.23
Appendix
Table 8: Full-model candidate accuracy at 0K background, using 100 examples per task and model. These are the baselines for the loss curves in Figure 20 .
Figure 20: Multi-fact loss sensitivity must be interpreted with baseline competence. Mean candidate-answer-loss degradation at 0K background. Color denotes task, and solid and dashed curves denote attention and contribution ranking. Each curve includes the full 100-example cohort, with zero degradation when N is at least the prompt length. Full-model accuracies are given in Table 8 .
5%
10%
Model
Ranking
Primary
Renorm
Primary
Renorm
Qwen-2.5-7B
attention
33
15
21
9
Qwen-2.5-7B
contribution
31
24
21
13
Gemma-7B
attention
226
7
206
4
Gemma-7B
contribution
227
11
205
7
Mistral-7B
attention
14
8
10
5
Appendix
Table 9: OpenWebText effective set-size estimates under deletion (Primary) and renormalization (Renorm) at 5% and 10% relative NLL tolerances. Both rankings are shown. Historical document pairing is unverified, so the comparison is at the corpus-summary level.
0.10 nat
0.20 nat
Model
Ranking
0K
1K
2K
4K
0K
1K
2K
4K
Qwen-2.5-7B
attention
16
–
–
16
8
–
–
16
Qwen-2.5-7B
contribution
32
–
–
32
8
–
–
16
Gemma-7B
attention
32
–
–
32
8
–
–
32
Gemma-7B
contribution
16
–
–
32
8
–
–
32
Mistral-7B
attention
8
–
–
16
8
–
–
8
Appendix
Table 10: Renormalized qa1 effective set-size estimates on the evaluated sizes at 0.10- and 0.20-nat mean answer-loss tolerances. Dashes indicate background conditions that were not evaluated for this control.
Figure 21: Renormalization changes both required set size and background sensitivity. (a) Ratios of renormalized to deletion estimates on OpenWebText from Table 9 . (b) Attention-ranked 4K/0K ratios at a 0.10-nat answer-loss tolerance. Mistral’s unresolved deletion result is marked at the tested endpoint and is not a resolved ratio. These controls are separate from the matched-target context-length evaluation.
Figure 22: Sensitivity to deletion varies substantially across layers and heads. Reference selected-set sizes are 128 for Gemma, 16 for Qwen, and eight for Mistral. All scanned layers and the ten largest measured Gemma head effects are shown. Layer and head evaluations use 50 and 20 documents, respectively. Separate interventions do not form an additive decomposition of the full-model requirement.
Figure 23: Loss-sensitive diagnostics improve exploratory prediction in this sample. Mean absolute log2 error in effective set size under leave-one-model-out evaluation. Lower is better. Labels report valid cases out of 16. Diagnostics with different unresolved-case patterns are not evaluated on identical sets of cases.
Figure 24: Positional and attention-ranked selection have different functional effects. The positional mask retains one initial source and N−1 recent sources. Both rules use deletion without renormalization at the same set sizes. At 2K and 4K, the positional control does not reach the 0.10-nat criterion within the tested grid. The experiment evaluates selection after computing dense attention.
Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes whether the model finds it. Gated attention cut first token attention from 46.7 percent to 4.8 percent at NeurIPS 2025, and Kimi K3 pairs that idea with Kimi Delta Attention and Attention Residuals behind a one million token window, eight times past the range where these diagnostics have been reported. This paper asks whether the fix survives that jump. We build SinkProbe, a suite that measures sink mass, massive activation, position resolved recall and the recency gap, and apply it to four small models that differ only in how they mix tokens and depth. Three results follow. The training objective produces the sink, not the architecture. Gating did not reproduce its published effect at our scale. Sink mass, activations and position bias moved independently. Code, data and the measurement protocol are released at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
Sara Rizwan, Samaanah Abdus Salam
Shadan Women’s College of Engineering and Technology Hyderabad, Telangana, India
Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on long-context data can help but is expensive due to the quadratic scaling of Attention. We observe that most tokens do not require (Global) Attention over the entire sequence and can rely on local context. Based on this, we propose L2A (Learning To Attend), a layer that enables conditional (token-wise) long-range memory access by deciding when to invoke global attention. We evaluate L2A on Qwen 2.5 and Qwen 3 models, extending their effective context length from 32K to 128K tokens. L2A matches the performance of standard long-context training to within 3% while skipping Global Attention for ∼80% of tokens, outperforming prior baselines. We also design custom Triton kernels to efficiently implement this token-wise conditional Attention on GPUs, achieving up to ∼2× improvements in training throughput and time-to-first-token over FlashAttention. Moreover, L2A enables post-training pruning of highly sparse Global Attention layers, reducing KV cache memory by up to 50% with negligible performance loss. Our code is released under Apache 2.0 at https://github.com/awslabs/hybrid-model-factory/tree/main/examples/research/L2A.
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.