Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61× speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.
Figures & tables
Figure 1: CacheRepair’s paradigm: a lightweight KV cache repair network.
Figure 2: Answer F1 versus p50 TTFT for Qwen2.5-14B on four downstream datasets.
Figure 3: Structured stale-KV error in Qwen2.5-3B on the first 256 MuSiQue requests. (a–b) Relative RMSE by layer and KV head. (c–d) Relative RMSE by layer and within-chunk position; outlines highlight boundary-local and layer-persistent patterns.
Figure 4: CacheRepair inference pipeline. Compressed stale-KV features and token embeddings condition the repair blocks. Each block receives stale-feature reinjection; the predicted residual is added to the full stale KV before global RoPE is applied to K.
Figure 5: Answer F1 versus p50 TTFT for Qwen2.5-3B on four downstream datasets.
Figure 6: Answer F1 versus p50 TTFT for Llama-3.1-8B on four downstream datasets.
Target LLM
H2D
Mask
Repair
RoPE
KV write
TTFT
Qwen2.5-3B
6.30
1.25
12.55
2.17
1.29
43.41
Llama-3.1-8B
20.78
1.37
18.14
6.56
3.63
77.85
Qwen2.5-14B
33.08
1.44
26.52
10.01
5.60
115.74
Table 1: Median online latency components (ms) for the largest repairers on MultiHop-RAG. TTFT also includes query processing, first-token generation, and runtime overhead. Component medians do not sum to median TTFT.
Variant
Params (M)
F1
EM
Δ F1 (pp)
TTFT p50 / p90 (ms)
Full Prefill
–
0.3155
0.232
–
68.24 / 90.81
Stale KV
–
0.1462
0.072
–
22.41 / 24.33
A0 CacheRepair
29.8
0.2462
0.172
0
28.86 / 33.23
A1 Token-causal
29.8
0.2421
0.172
−0.41
28.84 / 32.93
A2 Bidirectional
29.8
0.2427
0.164
−0.35
29.42 / 33.48
A3 Block diagonal
29.8
0.0133
0.000
−23.29
29.57 / 33.46
Table 2: Six repairer variants trained for two epochs and evaluated on the same 500 MuSiQue requests with Qwen2.5-3B. Δ F1 is computed from the reported mean F1 values relative to A0, in percentage points.
Target
Dataset
Repair
Baseline
Δ F1 (pp) and interval
Qwen2.5-3B
MuSiQue
30M
CacheBlend 0.1
+5.80 [+1.45, +10.15]
Qwen2.5-3B
HotpotQA
–
–
–
Qwen2.5-3B
MultiHop-RAG
51M
KV Packet
-1.53 [-5.65, +2.60]
Qwen2.5-3B
TriviaQA
51M
EPIC 32
-0.68 [-4.46, +3.11]
Llama-3.1-8B
MuSiQue
93M
CacheBlend 0.1
+6.69 [+2.07, +11.31]
Llama-3.1-8B
HotpotQA
93M
KV Packet
+10.21 [+5.49, +14.94]
Table 3: CacheRepair versus the highest-scoring baseline at half of Full Prefill p50 TTFT. Each method first selects its largest tested setting satisfying the budget, using latency alone. Brackets are simultaneous 95% paired-bootstrap bands across all eligible baseline methods within that panel; complete pairwise comparisons accompany the artifact. A dash indicates no eligible repairer.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Target LLM
Size
db
B
dseg
Parameters
Qwen2.5-3B
Small
256
5
8
9M
Medium
512
6
16
30M
Large
704
6
24
51M
Llama-3.1-8B
Small
256
5
8
23M
Medium
512
6
16
59M
Large
704
6
24
93M
Appendix
Table 4: Trained CacheRepair configurations for each target LLM. Parameter counts are rounded to the nearest million.
Dataset
Public pool
Examined
A
B
C
Union
Repeated
MuSiQue
2,417
2,375
705
1,695
264
1,818
7
HotpotQA
7,405
822
67
249
0
272
0
MultiHop-RAG
2,556
550
0
0
0
0
0
TriviaQA
17,944
1,123
540
379
1
541
32
Appendix
Table 5: Request selection for the four downstream datasets. Rule counts refer to examined candidates and can overlap; the union column counts requests excluded by at least one rule. Each dataset retains 500 evaluation requests and 50 reserves.
Target / dataset
Method
Full F1
Method F1
Δ F1 (pp)
W/L
Qwen2.5-3B MultiHop-RAG
CacheBlend 0.6
54.28
56.26
+1.98 [-0.023, +4.054]
19/10
Qwen2.5-3B MultiHop-RAG
KV Packet 8+8
54.28
55.09
+0.81 [-2.867, +4.447]
46/43
Llama-3.1-8B TriviaQA
CacheRepair 93M
80.49
82.30
+1.81 [-0.001, +3.655]
34/19
Llama-3.1-8B TriviaQA
InfoFlow 0.15
80.49
82.94
+2.45 [+0.556, +4.324]
41/17
Appendix
Table 6: Answer-level comparisons on the complete 500-request datasets. F1 is shown as a percentage. Differences are method minus Full Prefill, with pointwise 95% paired request-bootstrap intervals from 10,000 draws. W/L counts requests with higher/lower F1 than Full Prefill; the remaining requests have equal F1.
Dataset
CacheBlend 0.6
EPIC64
InfoFlow 0.25
Qwen2.5-3B / 51M
MuSiQue
+0.98 [-2.38, +4.35] 2.13 ×
-1.11 [-4.41, +2.17] 1.84 ×
+3.19 [-0.18, +6.56] 1.80 ×
HotpotQA
+3.74 [+0.63, +6.90] 1.72 ×
+3.83 [+0.65, +7.09] 1.30 ×
+4.43 [+1.19, +7.79] 1.79 ×
MultiHop-RAG
-2.69 [-5.00, -0.47] 3.67 ×
-1.05 [-3.66, +1.55] 2.14 ×
-0.50 [-3.00, +1.93] 2.38 ×
TriviaQA
-0.69 [-3.22, +1.92] 3.53 ×
+0.23 [-2.53, +3.00] 2.05 ×
+0.53 [-2.12, +3.20] 2.33 ×
Llama-3.1-8B / 93M
Appendix
Table 7: Largest repairer versus each highest tested recomputation preset on the same 500 requests per dataset. Each cell gives repair minus baseline F1 in percentage points, its pointwise 95% paired bootstrap interval, and baseline/repair p50 TTFT (speedup). Intervals are not adjusted jointly across the 36 comparisons.
MuSiQue
HotpotQA
MultiHop-RAG
TriviaQA
Method
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
Full Prefill
0.316 / 0.232
1.000
0.566 / 0.430
1.000
0.543 / 0.540
1.000
0.799 / 0.708
1.000
Stale KV
0.146 / 0.072
0.328
0.366 / 0.256
0.438
0.515 / 0.506
0.217
0.707 / 0.622
0.224
CacheBlend
0.198 / 0.120
0.463
–
–
0.534 / 0.528
0.416
0.773 / 0.694
0.415
EPIC
0.182 / 0.112
0.461
0.400 / 0.284
0.471
0.545 / 0.540
0.461
0.776 / 0.700
0.473
InfoFlow
–
–
–
–
0.539 / 0.532
0.496
0.774 / 0.690
0.489
Appendix
Table 8: Qwen2.5-3B quality at a TTFT budget of B=0.5Tfull . Each method uses its largest tested setting that meets the budget. Cells report F1/EM and p50 TTFT normalized by Full Prefill. Full Prefill and Stale KV provide reference values; a dash indicates no eligible setting.
MuSiQue
HotpotQA
MultiHop-RAG
TriviaQA
Method
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
Full Prefill
0.395 / 0.284
1.000
0.680 / 0.510
1.000
0.683 / 0.660
1.000
0.805 / 0.692
1.000
Stale KV
0.205 / 0.126
0.263
0.439 / 0.312
0.331
0.453 / 0.442
0.208
0.735 / 0.664
0.212
CacheBlend
0.285 / 0.204
0.494
–
–
0.643 / 0.630
0.494
0.820 / 0.734
0.473
EPIC
0.266 / 0.186
0.491
0.492 / 0.362
0.478
0.603 / 0.592
0.428
0.813 / 0.724
0.388
InfoFlow
–
–
–
–
0.627 / 0.614
0.498
0.815 / 0.730
0.458
Appendix
Table 9: Llama-3.1-8B quality at a TTFT budget of B=0.5Tfull . Each method uses its largest tested setting that meets the budget. Cells report F1/EM and p50 TTFT normalized by Full Prefill. Full Prefill and Stale KV provide reference values; a dash indicates no eligible setting.
MuSiQue
HotpotQA
MultiHop-RAG
TriviaQA
Method
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
F1 / EM
TTFT / Full
Full Prefill
0.450 / 0.348
1.000
0.733 / 0.580
1.000
0.731 / 0.722
1.000
0.837 / 0.740
1.000
Stale KV
0.176 / 0.120
0.202
0.388 / 0.288
0.247
0.675 / 0.672
0.165
0.774 / 0.686
0.172
CacheBlend
0.277 / 0.218
0.407
0.461 / 0.348
0.395
0.696 / 0.686
0.408
0.794 / 0.698
0.410
EPIC
0.284 / 0.214
0.406
0.522 / 0.400
0.488
0.697 / 0.688
0.492
0.800 / 0.704
0.328
InfoFlow
0.338 / 0.254
0.489
0.498 / 0.376
0.480
0.714 / 0.708
0.461
0.808 / 0.716
0.466
Appendix
Table 10: Qwen2.5-14B quality at a TTFT budget of B=0.5Tfull . Each method uses its largest tested setting that meets the budget. Cells report F1/EM and p50 TTFT normalized by Full Prefill. Full Prefill and Stale KV provide reference values; a dash indicates no eligible setting.
Target
Dataset
Capacity
Full F1
Repair F1
Δ F1 (pp)
Speedup
Qwen2.5-3B
MuSiQue
51M
31.55
26.26
-5.29 [-8.51, -2.00]
1.96
Qwen2.5-3B
HotpotQA
51M
56.61
55.35
-1.26 [-3.91, +1.39]
1.69
Qwen2.5-3B
MultiHop-RAG
51M
54.28
53.57
-0.71 [-2.95, +1.44]
2.98
Qwen2.5-3B
TriviaQA
51M
79.85
76.94
-2.91 [-5.15, -0.73]
2.86
Llama-3.1-8B
MuSiQue
93M
39.48
35.18
-4.30 [-7.18, -1.39]
2.63
Llama-3.1-8B
HotpotQA
93M
67.98
61.16
-6.82 [-9.68, -4.08]
2.38
Appendix
Table 11: Answer quality and latency of the largest repairer for each target. F1 is in percent; differences use paired 95% request-bootstrap intervals on each complete 500-request dataset. Speedup is Full Prefill p50 TTFT divided by repair p50 TTFT.
Figure 7: KV error before and after epoch-6 repair on Qwen2.5-3B (500 MuSiQue requests). K and V occupy separate rows; columns show stale error, repaired error, and error by distance from chunk start. Heatmap scales are shared within each row. Shading on the profiles shows 95% request-bootstrap intervals.
Target LLM
Width
Encoder
Fusion
Reinjection
Backbone
Head
Other
Total
Qwen2.5-3B
256
0.44
0.66
0.33
3.29
4.74
0.00
9.46
Qwen2.5-3B
512
1.48
1.57
1.58
15.76
9.46
0.00
29.84
Qwen2.5-3B
704
2.88
2.43
2.98
29.78
12.99
0.00
51.07
Qwen2.5-14B
320
2.76
1.84
0.51
5.14
31.56
0.00
41.81
Qwen2.5-14B
640
9.45
4.10
2.46
24.62
63.01
0.00
103.64
Qwen2.5-14B
832
17.72
5.65
4.85
48.52
81.89
0.00
158.62
Appendix
Table 12: Trainable parameter counts by component (millions). Fusion combines token embeddings with cache features; reinjection supplies cache features to each repair block. Width is db .
Figure 8: Residual relative RMSE on the separate set of 128 validation requests. PCA uses a basis and mean from the fitting requests; the learned-head projection uses each checkpoint’s fixed affine output space. Actual repair uses the corresponding checkpoint’s predictions.
Configuration
F1
EM
Δ F1 (pp)
TTFT (ms, p50)
GPU h
Full Prefill
0.3155
0.2320
+6.93
68.24
–
Stale KV
0.1462
0.0720
-10.00
22.41
–
MSE, 6 epochs
0.2560
0.1700
+0.98
33.28
–
MSE, 2 epochs
0.2462
0.1720
+0.00
28.86
23.10
QK/V, 2 epochs
0.1984
0.1260
-4.78
28.47
26.49
Huber, 2 epochs
0.2354
0.1560
-1.08
28.05
23.29
Appendix
Table 13: Qwen2.5-3B answer quality on 500 MuSiQue requests with a 30M repairer. The three two-epoch objectives share the same initialization, architecture, examples, sample order, and 25,000 updates. They use exact SDPA at the actual request length. F1 differences are computed from the reported means relative to two-epoch MSE (A0 in Table 2 ). Full, Stale, and the main six-epoch repairer provide serving references from the main evaluation runtime.
Configuration
Normalized KV MSE
KL(Full ∥ candidate)
Attention rel. RMSE
Stale KV
0.949 [0.930, 0.969]
1.731 [1.536, 1.936]
0.598 [0.589, 0.607]
MSE, 6 epochs
0.478 [0.470, 0.488]
0.555 [0.449, 0.674]
0.328 [0.315, 0.340]
MSE, 2 epochs
0.502 [0.493, 0.512]
0.698 [0.576, 0.832]
0.353 [0.340, 0.367]
QK/V, 2 epochs
0.606 [0.595, 0.618]
0.916 [0.781, 1.056]
0.405 [0.390, 0.420]
Huber, 2 epochs
0.508 [0.499, 0.518]
0.685 [0.569, 0.813]
0.351 [0.338, 0.365]
Appendix
Table 14: Functional effects on the same 256 MuSiQue requests. KL uses teacher forcing along the Full Prefill continuation and is averaged per request. Attention error is measured after the output projection at the last query position and pools squared error across requests and layers. Brackets give 95% request-bootstrap intervals.
Figure 9: Optimization with MSE, QK/V, and Huber on the same two-epoch training sequence. Both panels report normalized KV MSE using the shared residual statistics; each point averages 250 updates. The horizontal axes show updates and measured training time.
Configuration
F1 (%)
EM (%)
Δ F1 (pp)
TTFT (ms)
End-to-end (ms)
Full Prefill
33.13
25.00
+6.57 [+2.74, +10.56]
71.25
118.08
Stale KV
13.52
6.64
-13.03 [-18.35, -7.80]
23.11
77.75
CacheRepair
26.56
17.97
+0.00 [+0.00, +0.00]
32.44
82.96
Repair K
18.81
14.06
-7.75 [-12.26, -3.20]
32.72
82.33
Repair V
0.00
0.00
-26.56 [-31.54, -21.83]
32.70
453.00
Keep first chunk
25.89
18.36
-0.66 [-2.14, +0.75]
32.80
82.24
Appendix
Table 15: Output choices for the Qwen2.5-3B 30M repairer on 256 MuSiQue requests. Repair K/V applies the correction to that component; Keep first chunk retains its stale KV. Each variant executes the full repair network. F1 differences are paired against CacheRepair, with 95% request-bootstrap intervals. Latencies are p50; end-to-end time includes generation of up to 32 tokens.
Variant
First chunk (K/V)
Boundary (K/V)
Interior (K/V)
Stale KV
0.007 / 0.027
0.310 / 0.604
0.182 / 0.339
A0 CacheRepair
0.011 / 0.030
0.148 / 0.469
0.095 / 0.290
A1 Token-causal
0.017 / 0.037
0.150 / 0.470
0.095 / 0.290
A2 Bidirectional
0.021 / 0.038
0.149 / 0.469
0.095 / 0.290
A3 Block diagonal
0.115 / 0.157
0.191 / 0.503
0.126 / 0.310
A4 Entrance-only
0.010 / 0.029
0.149 / 0.468
0.095 / 0.289
Appendix
Table 16: Relative RMSE of K/V for the two-epoch variants on MuSiQue. Boundaries contain the first eight tokens of later chunks; interiors contain the remaining tokens. Errors and reference magnitudes are pooled across requests and layers.
Target LLM
KV
First chunk
Boundary
Interior
Qwen2.5-3B
K
0.007 [0.007, 0.007]
0.309 [0.307, 0.310]
0.181 [0.179, 0.183]
Qwen2.5-3B
V
0.027 [0.026, 0.027]
0.603 [0.600, 0.607]
0.336 [0.332, 0.340]
Llama-3.1-8B
K
0.011 [0.011, 0.011]
0.572 [0.570, 0.574]
0.320 [0.317, 0.322]
Llama-3.1-8B
V
0.026 [0.025, 0.026]
0.728 [0.724, 0.732]
0.437 [0.432, 0.442]
Qwen2.5-14B
K
0.011 [0.010, 0.011]
0.552 [0.549, 0.554]
0.311 [0.306, 0.315]
Qwen2.5-14B
V
0.026 [0.025, 0.027]
0.713 [0.709, 0.718]
0.391 [0.385, 0.396]
Appendix
Table 17: Stale-cache relative RMSE on the same 256 MuSiQue requests. Boundary and interior refer to later chunks; the boundary is the first eight tokens. Brackets give 95% request-bootstrap intervals.
Figure 10: Stale KV error by distance from a later chunk’s start, on the same 256 MuSiQue requests. Curves pool layers and heads; shading gives 95% request-bootstrap intervals.
Figure 11: Stale KV relative RMSE by target-LLM layer and position within later chunks, using the same 256 MuSiQue requests. Columns show canonical K and native V, with shared color scales across models. The 16 position bins reveal both boundary hotspots and error bands extending through chunk interiors.
Dataset
512
2048
8192
Δ [95% CI]
CR
MuSiQue
19.05
19.79
21.90
+2.84 [-0.03, +5.69]
25.60
HotpotQA
39.42
41.35
40.89
+1.47 [-1.61, +4.42]
53.92
MultiHop-RAG
55.09
54.53
55.05
-0.04 [-2.50, +2.45]
53.58
TriviaQA
72.26
72.83
73.58
+1.32 [-1.27, +3.88]
78.26
Appendix
Table 18: KV Packet calibration budget on Qwen2.5-3B. F1 (%) uses the same 500 requests per dataset. Δ is 8192 minus 512, in percentage points, with pointwise 95% paired-bootstrap intervals. CR is the fixed 30M CacheRepair reference.
Figure 12: p50 TTFT across document lengths for the three target LLMs, with 16 chunks and all 18 main-evaluation configurations. Each point contains three measurements of each of 16 fixed requests. Darker curves within a method indicate larger repairers or higher recomputation budgets. Both axes are logarithmic.
Figure 13: Dense matrix MACs across document lengths, for the same configurations as Figure 12 . Counts include target-LLM passes, repair, scoring, and first-token output heads. G denotes 109 MACs; both axes are logarithmic.
Figure 14: Peak serving-worker memory across document lengths, including target-LLM weights, KV pools, and online temporary tensors. Each point is the maximum across 48 timed runs; each worker loads one repairer capacity. Styles follow Figure 12 .
Target LLM
Configuration
TTFT p50 (ms)
TMACs
Peak
Extra
Qwen2.5-3B
Full Prefill
640.58
85.51
7.76
1.26
Qwen2.5-3B
Stale KV
76.06
0.32
9.90
3.39
Qwen2.5-3B
CacheRepair 51M
186.82
3.42
11.12
4.62
Qwen2.5-3B
CacheBlend 0.1
335.45
11.05
9.90
3.39
Qwen2.5-3B
EPIC 16
104.46
1.65
10.46
3.96
Qwen2.5-3B
InfoFlow 0.1
399.21
11.37
10.46
3.96
Appendix
Table 19: System cost at 16K document tokens and 16 chunks (16 requests, three measurements each). We show the largest repairers and the fixed baseline presets used in the chunk experiment. TMAC denotes 1012 dense matrix MACs. Peak is the serving-worker allocation, including weights, KV pool, and temporary tensors; Extra is its increment above the paired idle Full Prefill worker. Both memory columns are in GiB.
4K
8K
16K
Method
F1
TTFT
F1
TTFT
F1
TTFT
Full Prefill
48.63
129.9
50.59
274.8
52.54
648.1
Stale KV
50.13
27.6
42.29
43.2
37.19
74.8
CacheRepair 30M
46.68
38.7
49.02
67.4
49.80
145.8
KV Packet 512
52.15
26.2
50.78
41.4
48.05
75.7
CacheBlend 0.6
49.41
161.1
50.59
456.3
49.41
1487.4
Appendix
Table 20: Quality and latency on the same 128 MultiHop-RAG requests with Qwen2.5-3B and nested retrieved contexts. All seven evaluated configurations are shown. TTFT is in ms (p50); F1 is in percent. The final row gives CacheRepair minus Full Prefill F1 with pointwise 95% paired request-bootstrap intervals at each length.
Figure 15: Chunk-count sensitivity for Qwen2.5-3B on the same 500 MuSiQue requests. Document tokens, order, and queries remain fixed. CacheRepair uses the 30M epoch-six checkpoint; baseline settings are specified above. Error bars are pointwise 95% request-bootstrap intervals. The dashed line is the shared Full Prefill reference.
Method
C=1
C=2
C=4
C=8
Full Prefill
7.45 / 130.6
8.96 / 219.3
10.42 / 312.5
11.35 / 521.9
Stale KV
12.43 / 29.5
19.48 / 50.3
31.61 / 65.7
44.61 / 96.7
CacheRepair 30M
11.35 / 52.2
16.90 / 74.2
24.25 / 105.4
30.70 / 173.9
KV Packet 512
13.16 / 27.5
19.71 / 47.2
32.69 / 59.0
46.80 / 85.0
Appendix
Table 21: Closed-loop serving on 256 requests (64 per downstream dataset), Qwen2.5-3B, one A800. Each cell reports throughput (requests/s) and p95 submit-to-first-token latency (ms). C is the number of outstanding requests. One complete workload run is measured per method and concurrency level.
Target LLM
Repairer
Training
Statistics
File
Mean saving
Reuse
(M params.)
(GPU h)
(GPU h)
(MiB)
(ms/request)
(M requests)
Qwen2.5-3B
9
65.9
8.92
36.2
53.3
5.06
Qwen2.5-3B
30
70.1
8.92
114.0
50.4
5.65
Qwen2.5-3B
51
69.4
8.92
195.0
48.6
5.80
Llama-3.1-8B
23
81.9
11.56
89.1
126.9
2.65
Llama-3.1-8B
59
84.5
11.56
225.0
124.9
2.77
Appendix
Table 22: One-time costs of the nine epoch-six repairers. Training time covers recorded launches through the epoch-six checkpoint. Statistics are built once per target LLM and shared across its capacities. Checkpoint sizes are FP32 deployment files. The last column gives a latency-equivalent reuse count, using mean TTFT savings over the equally weighted four downstream datasets.
Category
Band definition
Realized range
Rows
Row share
Document tokens (share)
Mean/row
Short
1–1,024
261–1,013
5,000
10.0%
3,532,712 (3.55%)
706.5
Medium
1,025–2,048
1,376–1,812
5,000
10.0%
8,080,516 (8.12%)
1,616.1
Long
2,049–4,096
2,131–2,269
40,000
80.0%
87,926,653 (88.33%)
2,198.2
Total
1–4,096
261–2,269
50,000
100.0%
99,539,881 (100.00%)
1,990.8
Appendix
Table 23: Training examples by document length. Counts exclude query tokens; shares use 50,000 examples and 99,539,881 document tokens.
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.
Ruoling Qi, Yirui Liu, Xuaner Wu +5
Shanghai Jiao Tong University · Institute of Artificial Intelligence, China Telecom (TeleAI) · State University of New York at Buffalo
Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the prefill stage a dominant serving cost. RAG cache fusion reduces this cost by reusing precomputed key-value (KV) caches for retrieved chunks and selectively recomputing tokens under the current prompt. Existing selectors, however, face a dilemma between quality and efficiency: fast query-agnostic or final-layer query-to-context selectors can miss request-relevant evidence, whereas full-view query-aware selectors require broad context and layer visibility before recomputation and therefore stall the layer-wise cache-fusion pipeline. We present QCFuse, a compressed-view query-aware selector for RAG cache fusion. QCFuse uses chunk-anchor query probing to condition user-query states on compact per-chunk anchors and critical-layer profiling to identify recomputation tokens without all-layer inspection. We implement QCFuse in SGLang and evaluate it on four open-weight LLMs across six datasets. QCFuse reaches full-prefill-level quality. At matched quality, QCFuse achieves an average prefill-time speedup of 1.7x over full prefill and 1.5x over ProphetKV, the strongest quality-preserving baseline.
Jianxin Yan, Wangze Ni, Zhenxin Li +8
Zhejiang University, Hangzhou, China · Ant Group, Shanghai, China · The Hong Kong Polytechnic University, Hong Kong, China +5
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflects our core mechanism: much like assembling small tokens (or "coins") to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner. Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context. Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget.
Gyuwan Kim, Cheoneum Park, Tao Yang
University of California, Santa Barbara · Hanbat National University