The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.
Figures & tables
Figure 1: Proposed mechanism for implicit relative position encoding. (a) Tokens begin as initially uncorrelated. (b) Local mixing correlates nearby residual states. (c) Learned query and key projections can read out this lag-dependent structure as recency-biased global NoPE logits. The shading intensity denotes correlation or logit strength, while line width denotes mixing or attention weight.
Figure 3: Residual cosine during training. Post-attention cosine profiles at the first, middle, and last SWA layers ( w=128 ), with color denoting the fraction of training completed on a logarithmic scale. This demonstrates that residual stream recency bias increases with training and depth.
Figure 4: Window dependence of residual cosine. (a) RoPE-in-SWA and (b) NoPE-in-SWA, with the top panels showing the cosine gap c(d)−c(4096) at the middle SWA layer (L13), and the bottom panels showing the cosine gap c(1)−c(4096) across depth (crosses mark SWA layers and circles global NoPE layers). With or without RoPE, smaller windows result in stronger recency biases across depth.
Figure 5: Validation loss across window sizes. (a) RoPE-in-SWA and (b) NoPE-in-SWA. Mean 350M DCLM token cross-entropy (nats) over the final 30% of training. With or without RoPE, smaller windows result in lower loss.
Figure 6: Module changes in cosine gap and floor. After-minus-before changes in cosine gap c(1)−c(4096) (top) and cosine floor c(4096) (bottom), averaged across SWA and full-attention blocks at w=128 . Residual-addition steps compare with the isolated branch; shaded block summaries compare the complete update with its incoming residual. Note that the effect of each module on the similarity structure is largely as predicted by our theory.
Figure 7: Global NoPE logits during training. Head-mean profiles at the first, middle, and last global layers ( w=128 ), with color denoting the fraction of training completed on a logarithmic scale. This shows that, through training, the global NoPE heads learn to select the residual stream recency bias and thereby develop a corresponding recency bias in their logits.
Figure 8: Global NoPE logit gaps. RoPE-in-SWA (top) and NoPE-in-SWA (bottom), showing the head-mean gap ℓˉ(1)−ℓˉ(4096) (shading spans the full headwise range and vertical axes are symmetric-log). Though the window dependence is less clear in RoPE-in-SWA, it appears the logit recency bias collapses for large windows with NoPE-in-SWA.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Residual cosine at 350M. Post-attention gap c(d)−c(4096) : (a) every SWA layer at the final w=128 checkpoint; (b) first, middle, and last SWA layers across windows 64–4096.
Figure 10: Cross-moment magnitude and readout alignment. Final 350M RoPE-in-SWA model ( w=128 ). Left: ∥BX∥F at post-block residual states; crosses mark SWA and circles global NoPE layers. Middle: cos(BX,I) at SWA layers. Right: cos(BZ,M) at model-scale post-LN1 inputs to global heads; small points show eight heads and connected circles their unweighted mean. Magnitude uses a logarithmic axis; alignments use linear axes.
Figure 11: Global NoPE logits at 350M. Head-mean gap ℓˉ(d)−ℓˉ(4096) : (a) every global layer at the final w=128 checkpoint; (b) first, middle, and last global layers across window sizes.
Figure 12: 120M TextbookChapters robustness and validation loss. (a) Attention-branch cosine at initialization. (b–f) Post-attention residual cosine during training, across depth and SWA windows, including RoPE- and NoPE-in-SWA gaps. (g–h) Validation loss over the final 30% of training on 4096-token sequences. (i–l) Global-NoPE head-mean logit profiles and gaps during training, across depth and windows. Profile gaps are relative to lag 4096; window sizes are 64–4096. Window comparisons do not isolate the contribution of recency to the loss.
Figure 13: Random-token control: 350M, RoPE within SWA. (a) Attention-branch cosine at initialization. (b–e) Post-attention residual cosine during training, across depth and windows, including the depthwise gap. (f–i) Global-NoPE head-mean logit profiles and gaps during training, across depth and windows. Profile gaps are relative to lag 4096; window sizes are 64–4096.
Figure 14: Random-token control: 350M, NoPE within SWA. (a) Attention-branch cosine at initialization. (b–e) Post-attention residual cosine during training, across depth and windows, including the depthwise gap. (f–i) Global-NoPE head-mean logit profiles and gaps during training, across depth and windows. Profile gaps are relative to lag 4096; window sizes are 64–4096.
Figure 15: Random-token control: 120M, RoPE within SWA. (a) Attention-branch cosine at initialization. (b–e) Post-attention residual cosine during training, across depth and windows, including the depthwise gap. (f–i) Global-NoPE head-mean logit profiles and gaps during training, across depth and windows. Profile gaps are relative to lag 4096; window sizes are 64–4096.
Figure 16: Random-token control: 120M, NoPE within SWA. (a) Attention-branch cosine at initialization. (b–e) Post-attention residual cosine during training, across depth and windows, including the depthwise gap. (f–i) Global-NoPE head-mean logit profiles and gaps during training, across depth and windows. Profile gaps are relative to lag 4096; window sizes are 64–4096.
Figure 17: KDA–NoPE natural-language extension. Columns compare 350M (left) and 120M (right). (a–b) Residual cosine gap profiles across depth; (c–d) lag-1 residual gaps; (e–f) global-NoPE head-mean logit gap profiles; (g–h) lag-1 head-mean logit gaps. All gaps are relative to lag 4096.
Figure 18: KDA–NoPE random-token extension. Columns compare 350M (left) and 120M (right). (a–b) Residual cosine gap profiles across depth; (c–d) lag-1 residual gaps; (e–f) global-NoPE head-mean logit gap profiles; (g–h) lag-1 head-mean logit gaps. All gaps are relative to lag 4096.