Authors: Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Wajih Hassan Raza, Atta Ul Asad, Young D. Kwon, Michal Valko, Dean F. Hougen
Organizations: University of Oklahoma · Stanford University · University of Houston · Lahore University of Management Sciences · University of Cambridge · Isara Labs
Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projected leave-one-out attention-output deletion cost and aggregated across sampled futures with optional trajectory weighting. The temporary continuations are discarded before final decoding, requiring no auxiliary model or training. Ablations isolate the mechanism: at B=128, a single response-side continuation recovers about 89% of the gain over the prompt-window control, while additional futures provide smaller improvements. At B=128, LORE-KV raises the LongBench average on Qwen2.5-14B from 45.49 to 48.24 (+2.75) and the 16K RULER average on Mistral-7B from 45.20 to 51.05 (+5.85). Gains diminish at larger cache budgets and coexist with task-level regressions. LORE-KV incurs 1.46-2.77x AnDPro's per-sample wall-clock time as a one-time compression overhead across six dense and hybrid-attention backbones.
Figures & tables
Figure 1: Monte Carlo future-query estimation: lookahead queries better match future decoding queries, with coverage saturating near M=4 .
Figure 2: Equal-accuracy KV-cache efficiency on LongBench (Mistral-7B).
Qwen3.5-9B
Gemma-4-12B-it
Method
B=128
256
512
1024
B=128
256
512
1024
Full Cache
52.97
27.27
CriticalKV
46.30
48.30
49.58
50.31
19.74
21.53
22.45
23.22
AnDPro
46.89
48.72
50.08
50.42
20.21
21.92
22.73
23.47
LORE-KV /E
48.49
50.05
51.12
51.79
20.85
22.63
23.69
24.44
LORE-KV /D
48.57
50.22
51.16
51.92
20.93
22.76
23.81
24.59
Table 2: LongBench on hybrid-attention backbones. Eviction is applied only to the KV-growing layers.
Figure 3: RULER accuracy averaged over 13 tasks at 4K, 8K, and 16K context lengths. Per-task results are given in Tables 7 and 10 .
Figure 4: Accuracy-memory trade-off on Llama-3.1-8B-Instruct. Higher and further left indicates a better trade-off.
Figure 5: Attention-output fidelity. Similarity between retained- and full-cache outputs on future queries across layers. Mistral-7B-Instruct-v0.2, B=128 .
Figure 6: Gain versus query mismatch. Per-example gains in oracle retention and output fidelity with increasing prompt-to-future mismatch. Mistral-7B-Instruct-v0.2, B=128 , M=4 .
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Dataset
Method
128
256
512
1024
Mistral
LongBench
AnDPro
7.8799
7.4885
7.4355
7.5338
CriticalKV
6.6867
6.3991
6.3982
6.5276
LORE-KV /E
17.5342
19.5496
19.6767
17.7930
RULER
AnDPro
6.2360
4.5775
5.9707
3.1746
CriticalKV
6.1907
4.1849
4.9899
2.4039
LORE-KV /E
13.6877
11.1720
14.5537
7.2315
Appendix
Table 3: Average inference time per sample (seconds). Lower is better. LORE-KV uses fixed lookahead compute, so non-monotonic latency across budgets, including at B=1024 , is consistent with runtime variation.
Figure 7: Latency-quality trade-off. (a) Stage-wise latency averaged over two seeds and (b) quality versus total latency across cache budgets on Mistral-7B-Instruct-v0.2, averaged over LongBench, RULER, and NIAH.
Symbol
Meaning
Default
(a) Core method
w
prompt-end observation-window size
32
M
number of lookahead particles
4
T
lookahead particle length
8
λL
lookahead mixing coefficient in Eq. 7
1.0
ϵ
renormalization guard in Eq. 5
10−6
Appendix
Table 4: LORE-KV hyperparameters and default settings.
Task
Abbr.
Task type
Metric
Avg. len
Language
#Samples
NarrativeQA
NQA
Single-Doc. QA
F1
18,409
EN
200
Qasper
Qasp
Single-Doc. QA
F1
3,619
EN
200
MultiFieldQA-en
MF
Single-Doc. QA
F1
4,559
EN
150
HotpotQA
HQA
Multi-Doc. QA
F1
9,151
EN
200
2WikiMultihopQA
2WQA
Multi-Doc. QA
F1
4,887
EN
200
MuSiQue
Musq
Multi-Doc. QA
F1
11,214
EN
200
Appendix
Table 5: LongBench English datasets, the abbreviations used in all LongBench results tables, metrics, average input lengths, and sample counts ( Bai et al., 2024 ) .
Table 13: Lookahead geometry and compute scaling. LongBench on Mistral-7B-Instruct-v0.2 at B=128 : (a) LORE-KV /E across trajectory count M and horizon T ; (b) quality frontier under a matched lookahead-token budget versus AnDPro.
Axis
Variant
B=128
B=256
B=512
B=1024
Core
LORE-KV /E core
40.28
41.38
41.86
42.20
LORE-KV /D core
40.21
41.23
41.68
42.39
Lookahead budget
M=0 (prompt window only; no lookahead term)
35.87
37.50
38.70
39.50
M=1 (single surrogate response)
39.81
40.85
41.56
42.03
M=2
40.17
40.18
41.10
41.22
M=8
40.69
40.97
41.52
41.65
Appendix
Table 14: LongBench ablations (16-dataset average) with Mistral-7B-Instruct-v0.2.
B=128
B=256
M
Reliability weighting
Head budgets
Selection
4K
8K
16K
4K
8K
16K
AnDPro (strongest baseline)
54.85
50.06
45.20
71.05
64.08
57.91
1
unit weight
uniform
top- K
56.90
54.29
52.15
68.45
62.77
57.70
adaptive
top- K
56.92
54.31
52.18
68.50
62.81
57.75
adaptive
diversity
56.95
54.34
52.20
68.55
62.85
57.78
4
Log-probability only
uniform
top- K
58.39
55.72
50.98
70.08
64.45
58.85
Appendix
Table 15: RULER ablations with Mistral-7B-Instruct-v0.2 (13-task averages).
Figure 8: NIAH retrieval. Exact-match and word-recall scores across needle depths and 4K–32K contexts on Llama-3.1-8B-Instruct at B=1024 . Panel means are shown above each plot.
Figure 9: Retention-order shift. Candidate-score profiles, top- B overlap, and Spearman correlation across eviction criteria. Mistral-7B-Instruct-v0.2, B=128 , M=4 .
Figure 10: Criterion comparison. (a,d) Recall@ B against the future-query oracle. (b,e) Attention-output error (lower is better). (c,f) Layer-wise Spearman correlation with the oracle.
Figure 11: Retention by token class. Mean retained fraction (left) and class enrichment P(c∣R)/P(c∣X) (right).
Figure 12: Future-query signal across layers and heads. Layer perturbation energy (left) and retention-recall gain from multi-lookahead queries (right). Mistral-7B-Instruct-v0.2, B=128 , M=4 .
LORE-KV /E budget
128
256
512
1024
Ada-KV
∼ 332
∼ 745
> 1024
> 1024
CriticalKV
∼ 307
∼ 709
> 1024
> 1024
AnDPro
∼ 208
∼ 414
∼ 743
> 1024
Saving (%)
38.5
38.1
31.1
>0
Appendix
Table 16: Baseline budgets matching LORE-KV /E on LongBench (Mistral-7B-Instruct-v0.2).
Figure 13: Budget equivalence under behavioral fidelity. Fidelity across cache budgets (left) and baseline cache savings required to match LORE-KV /E (right).
Mehta Family School of Data Science and Artificial Intelligence, Indian Institute of Technology Roorkee · Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi · Department of Electrical Engineering, Indian Institute of Technology Delhi