cs.CLSep 28, 2026

RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings

Authors: Jarod Lévy, Mathurin Videau, Jad Yehya, Jean-Rémi King, Stéphane d'Ascoli, Thomas Moreau

Organizations: Meta AI, Paris · Inria, Université Paris-Saclay, Palaiseau, France

Abstract

Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LeRoPE: Learnable RoPE Frequencies Improve Language Modeling

    Jul 11, 2026Petros Karypis, Sean O'Brien, Shreyas Kadekodi +2Rotary Position EmbeddingPositional Encoding

  2. RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably

    May 15, 2026Yufeng Du, Phillip Harris, Minyang Tian +5Rotary Position EmbeddingEfficient Long-Context Inference