Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.
Figures & tables
Method
Slow-band geometry
Data-aware
Extrapolates
No in-domain
Throughput
position
(train-free)
penalty
(ktokens/s)
RoPE ( Su et al., 2024 )
Fixed token index
×
×
✓
✓ 61.7
NoPE ( Kazemnejad et al., 2023 )
None
×
∼
×
✓ 62.4
HoPE ( Chen et al., 2025 )
Removed
×
✓
✓
✓ 62.3
YaRN ( Peng et al., 2024 )
Rescaled token index
×
✓
×
✓ 62.2
CoPE ( Golovneva et al., 2024 )
Pairwise data-aware coordinate
✓
∼
✓
× 4.3
Table 1: Comparison of positional-encoding design properties. “Data-aware” indicates that the positional mechanism depends on the input. Throughput measures prompt-prefill speed on the 1B architecture at context 4096 on one A100-80GB (higher is better), using the median of nine timed iterations after warmup. ✓ yes, ∼ partial or with caveat, × no.
Fast bands angle
Slow bands angle
RoPE
mθ
mθ
DaRoPE
mθ
ch(xm)θ
Method 1: DaRoPE pseudocode and slow-band geometry. Left: the DaRoPE rotary code update. Right: fast- and slow-band coordinates under RoPE and DaRoPE, with the schematic example the cat is a feline . RoPE preserves token order, HoPE removes slow-band rotation, and DaRoPE places cat and feline nearby in its data-aware coordinate.
Figure 1: Language Models from 124M to 1B: In-Domain and Extrapolation Performance. Per-position perplexity on N=256 held-out PG-19 documents. Models train at context 4096 (dotted line) and evaluate to 16k without further training; YaRN is applied to RoPE-500k at inference time. Curves use a 128-token moving average; the y-axis is clipped at perplexity 70 for readability.
Commonsense and reasoning
Knowledge and language
Math and code
Overall
Method
HSwag
ARC-e
ARC-c
PIQA
OBQA
Wino
CSQA
COPA
MMLU
RACE
TQA
BBH
GSM8K
HEval
MBPP
Avg.
DaRoPE
72.9
72.9
52.1
77.5
43.2
67.9
55.1
85.0
45.7
39.9
55.8
34.1
29.9
37.8
50.6
54.7
HoPE
73.6
71.1
50.8
78.0
41.6
66.1
56.1
82.0
47.2
40.2
56.1
34.3
29.0
37.2
51.6
54.3
Table 2: In-context evaluation of the 50B-parameter MoE models. Higher is better; bold marks the best accuracy in each column.
LongBench, english only subset, 32k
CrossCodeEval, 4k
RepoBench, 32k
Method
Single QA
Multi QA
Sum.
Synth.
Code
Few-shot
Avg.
Edit
EM
NLL
Edit
EM
NLL
DaRoPE
17.4
21.1
12.0
1.8
57.1
58.4
27.7
60.3
11.6
1.82
62.6
31.7
2.90
HoPE
17.0
14.6
12.5
3.0
56.6
55.9
26.1
60.3
11.4
1.83
57.6
28.6
3.22
Table 3: Long-context evaluation of the 50B MoE models. LongBench ( Bai et al., 2024 ) uses official metrics averaged within category; CrossCodeEval ( Ding et al., 2023 ) and RepoBench ( Liu et al., 2024b ) report edit similarity, exact match, and answer NLL. Higher is better except for NLL; bold marks the best value in each column. Each method is represented by one trained model.
Figure 2: Retrieval through model depth. Rows: return-from-digression (top) and key–value recall (bottom). ( left ): accuracy across difficulty ( N=800 ; 95% binomial CIs); return compares the correct answer with the recent decoy, while key–value uses exact top-1 selection. ( center ): At the hardest difficulty, correct token’s log-probability advantage ( N=200 ; 95% CIs; gray marks the final ten) for each layer through the final normalization and output head. ( right ): log attention ratio =log[a(correct)/a(competitor)] averaged over all 16 heads; red favors correct and blue the recent decoy (digression) or mean of the other K−1 values (key-value). No head or layer is selected.
Figure 3: Test NLL on four non-textual datasets. Bars show means across three seeds; error bars are 95% within-example confidence intervals. The right panel reports mean rank across the four datasets, with standard errors across datasets. Holm-corrected paired tests use test examples across three seeds. Significance is relative to DaRoPE: ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 . Test sizes are N=77 (JSB), 639 (MAESTRO), 8399 (HRG), and 57,782 (Sleep-EDF). NoPE is off-axis: 1.1634 , 1.5075 , 4.3494 , and 4.4320 , respectively.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
124M
350M
1B
dmodel
768
1024
2048
layers
12
24
25
heads
12
16
16
head dim d
64
64
128
fast bands K=d/4
16
16
32
context L
4096
4096
4096
Appendix
Table 4: Per-scale language-model architecture and training. Head dimension is dmodel/nheads ; the wavelength split θj<2π/L yields K=d/4 fast bands at θ=500k,L=4096 .
Variant
In-domain NLL ↓
8–16k NLL ↓
DaRoPE (per-head α )
2.9447
3.957
No β
2.9420
4.066
Fixed α
2.9437
4.449
Per-band α
2.9439
4.815
Dynamic α
2.9451
4.281
Appendix
Table 5: DaRoPE formulation ablation at 124M. One training seed; held-out in-domain and 8–16k NLL, lower is better.
Method
124M
350M
1B
RoPE-10k
18.95
15.16
17.70
RoPE-500k
18.92
15.10
17.67
NoPE
250.63
20.15
19.56
HoPE
18.94
15.13
17.68
PoPE
19.22
15.28
17.81
DaRoPE
18.95
15.09
17.70
Appendix
Table 6: In-domain validation perplexity at the three dense scales. Perplexity (lower is better) at the training length L=4096 on a held-out split of each model’s own pretraining distribution: FineWeb-Edu at 124M/350M ( 1.94 M tokens) and DCLM at 1B ( 4.01 M tokens). Values are comparable within a column, not across columns.
Figure 5: Per-position validation negative log-likelihood . The loss is averaged over N=256 documents for the HoPE and DaRoPE 50B models. Both models are trained with a context length of 4096 tokens, marked by the vertical dashed line, and evaluated on sequences of up to 16,384 tokens. Their curves nearly overlap both within and beyond the training window.
Figure 6: Controlled long-context performance across RULER and BABILong. Lines show task-level median accuracy and CE gain; shaded regions show the interquartile range. The dotted line marks the 4k training context. The two methods are mixed across tasks and context lengths.
Training hyperparameter
Value
Training steps
65,392
Sequence length
4,096
GPUs
256
Micro-batch / GPU
1 sequence
Gradient accumulation
9
Global batch size
2,304 sequences
Appendix
Table 7: Shared training hyperparameters for the HoPE and DaRoPE MoE models.
Architecture hyperparameter
Value
Total parameters
49.49B
Active parameters / token †
2.57B
Transformer layers
28
Model dimension
2,560
Attention heads
20
Head dimension
128
Appendix
Table 8: Shared architecture hyperparameters for the MoE models. Fine-grained routed experts with a shared expert follow DeepSeekMoE ( Dai et al., 2024 ) ; sigmoid routing with bias-based load balancing follows DeepSeek-V3 ( DeepSeek-AI et al., 2024 ; Wang et al., 2024a ) .
Figure 7: Attention ratios for the remaining methods and coordinate intervention. Rows show return-from-digression at D=2048 (top) and key–value recall at K=256 (bottom), matching Figure 2 . Columns show RoPE-500k, PoPE, HoPE, NoPE, and the trained DaRoPE model after its coordinates are shuffled across token positions. Each cell averages all 16 heads in that layer. Red favors the correct value; blue favors the recent decoy or mean competitor. All panels use the same N=200 prompts, row order, and color scale; the dashed line precedes layers 15–24.
Figure 8: Full synthetic-task results. Per-task token accuracy (mean ± SD across seeds) in domain and at 2× and 8× the training length, with task-family and aggregate summaries.
Task
Example
Answer
⟨ ch ⟩
Description
Regular
even_pairs
EP:01110011 → 1
single
50
Is the number of adjacent differing pairs even?
parity
P:01110011 → 1
single
50
Parity of the number of 1 s.
cycle_navigation
CY:11220022 → 2
single
20
Position on a 5-cycle after a walk ( 0 / 1 / 2 = left/stay/right).
modular_arithmetic
MA:2-3*0+4 → 1
single
20
Evaluate a flat expression mod 5 (the paper’s simple variant; no brackets).
Deterministic context-free
Appendix
Table 9: Chomsky-hierarchy tasks. × excluded from our results: at depth 4 with a 256-token training context, exact-match is 0 for every encoding we ran them with, so no instance is solved in this configuration .
Task
Example
Answer
⟨ ch ⟩
Description
needle
N:45790189 → 00000001
sequence
6
At each step, has the current symbol occurred earlier?
induction
I:55791289 → 05000001
sequence
2
Given … AB … A , predict B (induction head).
Appendix
Table 10: Retrieval probes. Pure content matching: the answer never depends on absolute position.
Task
Example
Answer
⟨ ch ⟩
Description
fixed_offset
FO:3141592#0003 → 5
single
10
Pure position: return the digit k places from the end.
dup_key_recall
DK:a3b7a9c1?a → 9
single
10
Position + content: the query key occurs twice; return the later value.
assoc_recall
AR:a3b7c1?b → 7
single
10
Content only: the query key occurs once; return its value.
Appendix
Table 11: Positional probes. assoc_recall is the control: it is identical to dup_key_recall except that the query key is unique, so any gap between the two is attributable to the positional requirement rather than to capacity.