Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-frequency components support positional sensitivity while potentially disrupting semantic stability. We also derive a theoretical context-length bound beyond which, under specified conditions, a fixed attention-score comparison cannot jointly avoid semantic reversal and positional insensitivity. Guided by our fresh theoretical insights, we introduce RoPE Profiler, a lightweight, plug-and-play diagnostic toolkit that augments existing evaluations with zero additional forward passes by reusing cached query and key activations. Reusing activations collected during evaluation, the toolkit incurs little overhead. It supplements standard benchmark scores with two diagnostic scores that reveal semantic and positional weaknesses and help users prioritize which aspect to address. Crucially, our evaluations across 49 long-context task settings reveal a distinct pattern where reasoning tasks predominantly suffer from semantic reversal, whereas retrieval tasks are primarily vulnerable to positional insensitivity. Guided by our theory and diagnostic profiles, targeted high-frequency rescaling achieves immediate gains without additional training, improving task accuracy by up to 20 percentage points on Qwen3-8B and 25 percentage points on Llama-3.1-8B-Instruct.
Figures & tables
Figure 1 : Illustrative task consequences of two RoPE failure modes. 1(a) For query country , the scores of keys France and chair reverse order between virtual relative distances 0 and m⋆ . As attention directs information routing toward higher scores, this reversal favors the distractor chair over the semantically relevant France at relative distance m⋆ and can contribute to an incorrect prediction. 1(b) For query A and key B , the vertical bracket marks a marginal normalized score gap between adjacent virtual relative distances, gz(m)<ζ . This leaves the model unable to distinguish between the two adjacent positions.
Figure 2 : (a) Under our high-frequency split, the predicted score distribution closely matches the empirical distribution. (b,c) Defining the high-frequency band. The curves show how randomness changes as more components are included. Our theory-guided split selects the largest band satisfying the split criterion while retaining high randomness. The trends vary with context length, and larger M allows more high-frequency components, consistent with our theory. (See Appendix D.2 .)
Figure 3 : Augmenting accuracy-based evaluation with failure diagnostics. For a given model and the 49 reference task configurations (described in § E ), we use query and key activations from standard evaluation to compute relative semantic failure susceptibility wS . A higher wS suggests greater susceptibility to semantic reversal relative to positional insensitivity within the reference set.
Figure 4 : Accuracy gains from our exploratory intervention search. Bars show the best observed intervention in each setting’s assigned semantic or positional direction minus the original baseline on identical examples within each setting. The baseline is excluded from the candidate maximum. Each model’s 49 task/configuration settings (described in Appendix E ) are ordered from left to right by increasing semantic failure susceptibility, from blue to orange.
Figure 5 : Semantic failure susceptibility across tasks. Points compare ranks of semantic failure susceptibility wS in Qwen3-8B and Llama-3.1-8B-Instruct for the 49 task/configuration settings described in Appendix E . Most tasks show similar patterns across models, while three adjacent-element retrieval tasks differ.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Pointwise consistency and tightness audit of the sharper coefficient-specific local positional-response envelope and local-response floor from the appendix. Panels (a)–(c) form the Qwen3-8B row and panels (d)–(f) form the Llama-3.1-8B row; every point is one first-layer query–head–key coefficient vector. Left: the observed integer-window maximum gap ζi against the coefficient-specific upper certificate T(ai) on logarithmic axes. Middle: the observed gap (blue) and upper certificate (orange) against rH(ai;M) ; solid curves and shaded regions are medians and 10 – 90% ranges over 32 equal-count rH bins, not an rH -only predictor. Right: the pointwise lower bound fpos(ai;M,ζi) against the observed certified-high norm share. Dashed diagonals indicate equality in the left and right columns. The Qwen3 row is an in-assumption geometric-grid audit, whereas the non-geometric Llama row is an out-of-assumption stress test. The right column uses the same point to set ζi and therefore measures consistency and tightness rather than out-of-sample predictive accuracy.
Task
What the model must do
Configurations
K
N
RULER2 (NeMo Skills)
Multi-key (MK)
Retrieve by key; harder variants answer retrieved MMLU questions.
Basic, easy, medium, hard
4
100
Multi-value (MV)
Retrieve multiple values or select an ordered question for a key.
Basic, easy, medium, hard
4
100
Question answering
Retrieve supporting documents and answer HotpotQA questions.
Basic, easy, medium, hard
4
100
LongBench-v1 ( Bai et al., 2024 )
Qasper
Answer questions about an academic paper.
Native chat ≤16 k
1
25
Appendix
Table 1 : The complete 49-setting diagnostic reference set. Each configuration in a row is a separate setting. K is the number of settings and N is the number of examples per setting and per model. Lengths are nominal benchmark configurations; the diagnostic uses each model’s actual input length.
Figure 7 : Head selection from the 49 reference settings. Each cell shows the sample variance of a head’s setting-level mean Common-pair rH , using a shared color scale across models. Black outlines identify the top-5% heads used by the main diagnostic and first intervention stage. Orange outlines identify the additional heads included in the top-10% set for the second stage. Layer and query-head indices start at one. All settings within a model use the same selected heads.
Setting
wS (%)
n/N
Base
NR
0
0.5
QA (basic)
54.2
100/100
78.00
70.00 (-8.00)
76.00 (-2.00)
75.00 (-3.00)
QA (easy)
54.2
100/100
76.00
66.00 (-10.00)
70.00 (-6.00)
77.00 (+1.00)
Multi-hop QA (up to 16k)
53.1
25/25
28.00
24.00 (-4.00)
24.00 (-4.00)
24.00 (-4.00)
Rule deduction (context size 3,000)
51.0
25/25
72.00
88.00 (+16.00)
88.00 (+16.00)
76.00 (+4.00)
Person location (16k)
61.5
25/25
76.00
80.00 (+4.00)
80.00 (+4.00)
68.00 (-8.00)
Previous location (8k)
56.2
25/25
28.00
32.00 (+4.00)
28.00 (+0.00)
32.00 (+4.00)
Appendix
Table 2 : Qwen3-8B: semantic direction on the top-5% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points).
Setting
n/N
Base
NR
0.25
0.5
0.75
QA (basic) †
100/100
78.00
63.00 (-15.00)
82.00 (+4.00)
80.00 (+2.00)
78.00 (+0.00)
QA (easy)
—
—
—
—
—
—
Multi-hop QA (up to 16k)
25/25
28.00
24.00 (-4.00)
24.00 (-4.00)
24.00 (-4.00)
28.00 (+0.00)
Rule deduction (context size 3,000)
—
—
—
—
—
—
Person location (16k)
—
—
—
—
—
—
Previous location (8k)
—
—
—
—
—
—
Appendix
Table 3 : Qwen3-8B: semantic direction on the top-10% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points). Dashes mark settings without a second-stage evaluation; † marks a generation limit that differs from the saved baseline.
Setting
wS (%)
n/N
Base
1.5
2
2.5
MK (basic)
49.0
100/100
99.00
100.00 (+1.00)
100.00 (+1.00)
100.00 (+1.00)
MK (easy)
16.7
100/100
96.00
97.00 (+1.00)
97.00 (+1.00)
97.00 (+1.00)
MK (medium)
22.9
95/100
72.63
72.63 (+0.00)
70.53 (-2.11)
70.53 (-2.11)
MK (hard)
32.3
95/100
64.21
67.37 (+3.16)
68.42 (+4.21)
70.53 (+6.32)
MV (basic)
46.9
100/100
22.00
15.00 (-7.00)
19.00 (-3.00)
20.00 (-2.00)
MV (easy)
15.6
94/100
43.62
43.62 (+0.00)
42.55 (-1.06)
37.23 (-6.38)
Appendix
Table 4 : Qwen3-8B: positional direction on the top-5% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points).
Setting
n/N
Base
1.25
1.5
1.75
2
MK (basic)
—
—
—
—
—
—
MK (easy)
—
—
—
—
—
—
MK (medium) †
95/100
72.63
70.53 (-2.11)
69.47 (-3.16)
67.37 (-5.26)
67.37 (-5.26)
MK (hard)
—
—
—
—
—
—
MV (basic) †
100/100
22.00
22.00 (+0.00)
24.00 (+2.00)
27.00 (+5.00)
28.00 (+6.00)
MV (easy)
95/100
44.21
40.00 (-4.21)
37.89 (-6.32)
33.68 (-10.53)
28.42 (-15.79)
Appendix
Table 5 : Qwen3-8B: positional direction on the top-10% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points). Dashes mark settings without a second-stage evaluation; † marks a generation limit that differs from the saved baseline.
Setting
wS (%)
n/N
Base
NR
0
0.5
QA (easy)
64.6
80/100
80.00
86.25 (+6.25)
87.50 (+7.50)
86.25 (+6.25)
Multi-hop QA (up to 16k)
67.7
25/25
24.00
20.00 (-4.00)
16.00 (-8.00)
16.00 (-8.00)
Person location (16k)
77.1
25/25
76.00
76.00 (+0.00)
84.00 (+8.00)
84.00 (+8.00)
Previous location (8k)
62.5
25/25
28.00
36.00 (+8.00)
28.00 (+0.00)
32.00 (+4.00)
Previous location (16k)
80.2
25/25
20.00
28.00 (+8.00)
24.00 (+4.00)
24.00 (+4.00)
Directional relations (8k)
69.8
25/25
36.00
40.00 (+4.00)
36.00 (+0.00)
40.00 (+4.00)
Appendix
Table 6 : Llama-3.1-8B-Instruct: semantic direction on the top-5% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points).
Setting
n/N
Base
NR
0.25
0.5
0.75
QA (easy)
—
—
—
—
—
—
Multi-hop QA (up to 16k)
25/25
24.00
24.00 (+0.00)
8.00 (-16.00)
8.00 (-16.00)
16.00 (-8.00)
Person location (16k)
—
—
—
—
—
—
Previous location (8k)
—
—
—
—
—
—
Previous location (16k)
—
—
—
—
—
—
Directional relations (8k)
—
—
—
—
—
—
Appendix
Table 7 : Llama-3.1-8B-Instruct: semantic direction on the top-10% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points). Dashes mark settings without a second-stage evaluation; † marks a generation limit that differs from the saved baseline.
Setting
wS (%)
n/N
Base
1.5
2
2.5
MK (basic)
55.2
81/100
98.77
100.00 (+1.23)
97.53 (-1.23)
27.16 (-71.60)
MK (easy)
11.5
100/100
83.00
80.00 (-3.00)
5.00 (-78.00)
0.00 (-83.00)
MK (medium)
8.3
100/100
61.00
61.00 (+0.00)
24.00 (-37.00)
1.00 (-60.00)
MK (hard)
20.8
81/100
54.32
48.15 (-6.17)
1.23 (-53.09)
0.00 (-54.32)
MV (basic)
54.2
100/100
34.00
20.00 (-14.00)
7.00 (-27.00)
0.00 (-34.00)
MV (easy)
14.6
81/100
27.16
19.75 (-7.41)
0.00 (-27.16)
0.00 (-27.16)
Appendix
Table 8 : Llama-3.1-8B-Instruct: positional direction on the top-5% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points).
Setting
n/N
Base
1.25
1.5
1.75
2
MK (basic)
—
—
—
—
—
—
MK (easy) †
100/100
83.00
79.00 (-4.00)
71.00 (-12.00)
68.00 (-15.00)
61.00 (-22.00)
MK (medium) †
100/100
61.00
60.00 (-1.00)
57.00 (-4.00)
47.00 (-14.00)
31.00 (-30.00)
MK (hard) †
81/100
54.32
44.44 (-9.88)
40.74 (-13.58)
23.46 (-30.86)
11.11 (-43.21)
MV (basic) †
100/100
34.00
19.00 (-15.00)
5.00 (-29.00)
1.00 (-33.00)
6.00 (-28.00)
MV (easy)
81/100
27.16
29.63 (+2.47)
30.86 (+3.70)
12.35 (-14.81)
0.00 (-27.16)
Appendix
Table 9 : Llama-3.1-8B-Instruct: positional direction on the top-10% heads. Entries give accuracy in percent and the change from the paired baseline in parentheses (percentage points). Dashes mark settings without a second-stage evaluation; † marks a generation limit that differs from the saved baseline.