Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.
Figures & tables
Method
Accuracy Preservation
Repeated Prefill Reduction
Existing Adapter Support
Selective Recomputation
✗
✓
✓
BaseShared
✓
✗
✓
BaseLRShared
✓
✓
✗
PReCache (Ours)
✓
✓
✓
Table 1: Conceptual comparison of KV cache sharing strategies for multi-LoRA agents across three criteria. ✓ and ✗ denote support and no support, respectively.
Figure 1: Overview of KV cache construction in PReCache across consecutive agent turns. (A1–A3) PreLRShared uses the hidden states generated when each context segment is first processed to construct the LR caches for all agents, eliminating the separate full-length backbone processing over the accumulated context LP required by BaseShared. (B1–B5) ReBaseShared reconstructs the shared base cache from adapter-free hidden states to reduce dependency on the previous agent’s adapted representation. ReBaseShared DB (B2) schedules the adapted and adapter-free paths together during both prefill and decoding, whereas ReBaseShared LP (B3–B4) performs adapter-free reconstruction as a contiguous prefill after the current agent’s turn.
Figure 2: Effect of reconstructing the shared base cache from adapter-free hidden states in Ministral-8B on a HotpotQA trajectory. (a) compares the previous agent’s and neutral base caches against the current agent’s base cache. (b) reports the layer-wise relative errors in the hidden states and base caches of PreLRShared and ReBaseShared. (c) reports the last-layer relative error across agent transitions and includes FullShared as a reference for entire KV cache reuse.
HotpotQA
ScienceQA
Model
Method
Easy
Medium
Hard
Avg.
1–4
5–8
9–12
Avg.
LLaMA-3.1-8B
NonShared
40.55
41.05
30.95
37.52 0.00
70.33
59.67
77.58
69.19 0.00
FullShared
38.05
37.90
26.15
34.03 − 3.48
67.71
56.42
72.75
65.62 − 3.57
DroidSpeak
39.80
38.30
27.60
35.23 − 2.28
68.17
59.04
74.92
67.37 − 1.82
CacheBlend
39.15
38.55
27.10
34.93 − 2.59
67.90
58.38
73.71
66.66 − 2.53
RelayCaching
40.15
39.55
29.20
36.30 − 1.22
68.79
59.25
75.83
67.96 − 1.23
Table 2: Mean benchmark accuracy (%) of NonShared and KV cache sharing methods on HotpotQA and ScienceQA. The value beside each average denotes its percentage-point difference from NonShared. For each backbone and benchmark, the higher and lower halves of the methods ranked by average accuracy are highlighted in green and red, respectively.
Figure 3: Single-stream TTFT ( ↓ ) and per-request throughput ( ↑ ) over trajectory length for both backbones on a single A6000 GPU. ReBaseShared uses lazy prefill scheduling.
Figure 4: Serving TTFT percentiles ( ↓ ) and per-request throughput ( ↑ ) over request rate on vLLM with a fixed 17.3 k-token trajectory on a single A100 GPU. ReBaseShared uses double batching.
Figure 5: Efficiency comparison of LP and DB scheduling for ReBaseShared. The left panels report single-stream TTFT ( ↓ ) and per-request throughput ( ↑ ) over trajectory length. The right panels report p50 TTFT ( ↓ ) and per-request throughput ( ↑ ) over request rate with a fixed 17.3 k-token trajectory under concurrent serving.
Figure 6: Peak GPU memory usage across trajectory lengths for LLaMA-3.1-8B.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Layer-wise geometry of neutral base cache reconstruction in ReBaseShared. Each quantity is computed per token and then averaged across tokens, prompts, and agent transitions. Solid and dashed lines represent base cache and hidden state measurements, respectively. At the token level, ρ>2cosθ , or equivalently Γerr>1 , indicates lower relative error for the neutral state than for the state constructed from the previous agent’s hidden states.
Method
TTFT (s)
Model Latency (s)
E2E Latency (s)
Trajectory Length (tokens)
Average
Maximum
NonShared
1.41
7.23
16.21
1268
6857
FullShared
0.90
9.20
20.64
1757
7078
DroidSpeak
1.29
7.10
16.11
1340
6317
CacheBlend
1.35
6.98
16.64
1569
7103
RelayCaching
1.31
8.00
17.58
1455
6982
Appendix
Table 3: Average benchmark latency and trajectory length for LLaMA-3.1-8B on HotpotQA. Latencies are reported in seconds, with average and maximum trajectory lengths reported in tokens.
Figure 8: Accuracy-latency trade-off for LLaMA-3.1-8B on HotpotQA. The panels compare benchmark accuracy with TTFT, model latency, and end-to-end latency, respectively. Higher accuracy and lower latency are preferred.
Benchmark
NonShared
FullShared
DroidSpeak
CacheBlend
RelayCaching
BaseShared
PreLRShared
ReBaseShared
HotpotQA
0.21
0.44
0.28
0.34
0.31
0.25
0.30
0.32
ScienceQA
0.28
0.60
0.34
0.35
0.39
0.35
0.39
0.37
Appendix
Table 4: Standard deviation of benchmark accuracy across 20 complete evaluation runs, reported in percentage points.
Model
Metric
Method
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
LLaMA-3.1-8B
TTFT (s)
NonShared
1.93
2.53
3.72
6.77
16.36
33.72
73.18
FullShared
1.14
1.27
1.63
2.39
4.22
9.05
23.15
DroidSpeak
1.62
2.15
3.21
5.55
11.12
25.23
67.82
CacheBlend
1.67
2.15
2.31
3.72
8.61
19.05
56.38
RelayCaching
1.87
2.35
2.44
4.67
9.89
22.07
71.19
BaseShared
1.61
2.13
3.08
5.25
10.54
23.80
68.19
Appendix
Table 5: Single-stream TTFT in seconds and throughput in tokens per second over trajectory length on a single A6000 GPU. ReBaseShared uses LP scheduling, and TP denotes throughput.
Metric
Method
QPS
0.5
1
2
4
8
16
TP (tok/s)
NonShared
1573
1282
596
70
52
48
FullShared
2922
2910
2818
1809
205
164
DroidSpeak
1592
1416
686
78
60
53
CacheBlend
1939
1847
1391
140
104
83
RelayCaching
1827
1678
835
93
68
59
Appendix
Table 6: Serving TTFT percentiles in seconds and per-request throughput in tokens per second over request rate with a fixed 17.3 k-token trajectory on a single A100 GPU. ReBaseShared uses DB scheduling, and TP denotes throughput.
Metric
Schedule
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
TTFT (s)
LP
1.23
1.36
1.94
3.19
6.06
12.53
27.24
DB
1.24
1.58
2.33
3.94
7.72
17.95
49.29
TP (tok/s)
LP
166
245
404
635
900
1087
1068
DB
126
189
299
486
723
907
889
Appendix
Table 7: LP and DB scheduling for ReBaseShared under single-stream inference on LLaMA-3.1-8B. TP denotes throughput.
Metric
Schedule
QPS
0.5
1
2
4
8
16
p50 TTFT (s)
LP
0.06
0.06
6.97
18.26
22.67
24.79
DB
0.07
0.07
0.21
1.29
4.84
6.83
TP (tok/s)
LP
1670
1371
947
69
47
38
DB
2036
1972
1408
182
103
85
Appendix
Table 8: LP and DB scheduling for ReBaseShared under concurrent serving with a fixed 17.3 k-token trajectory. TP denotes throughput.
Method
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
NonShared
15.69
16.07
16.83
18.33
21.35
27.39
39.47
FullShared
15.29
15.47
15.85
16.60
18.10
21.11
27.13
DroidSpeak
15.52
15.86
16.48
17.75
20.29
25.37
35.50
CacheBlend
15.49
15.77
16.36
17.52
19.86
24.52
33.86
RelayCaching
15.46
15.73
16.28
17.38
19.59
24.01
32.84
BaseShared
15.30
15.49
15.88
16.66
18.22
21.34
27.58
Appendix
Table 9: Peak GPU memory usage over trajectory length on LLaMA-3.1-8B.
Metric
Method
N=4
N=8
TTFT (s)
NonShared
68.36
90.05
FullShared
23.98
22.73
DroidSpeak
66.57
88.89
CacheBlend
48.71
72.13
RelayCaching
52.99
80.96
BaseShared
58.09
87.60
Appendix
Table 10: Efficiency and peak GPU memory usage for different numbers of agents on LLaMA-3.1-8B-Instruct using a single A6000 GPU. Each turn adds 8 k tokens over a total of eight turns.
Method
r=4
r=8
r=16
r=32
NonShared
34.72
36.72
36.70
36.80
RelayCaching
32.95
34.63
34.47
34.20
PreLRShared
33.51
34.77
34.67
34.53
ReBaseShared
34.41
36.08
35.98
35.81
Appendix
Table 11: HotpotQA accuracy (%) across LoRA ranks on Ministral-8B-Instruct.
Method
r=8
r=16
r=32
r=64
NonShared
768
762
756
750
FullShared
1693
1696
1694
1690
DroidSpeak
937
927
921
913
CacheBlend
1042
1049
1042
1033
RelayCaching
900
902
891
878
BaseShared
976
967
958
953
Appendix
Table 12: Single-stream throughput in tokens per second across LoRA ranks at a fixed 33.7 k-token trajectory on LLaMA-3.1-8B-Instruct.
Metric
Method
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
TTFT (s)
NonShared
1.81
4.16
4.39
6.45
12.91
27.01
68.73
FullShared
1.45
1.45
2.00
2.71
5.23
10.18
24.62
DroidSpeak
1.76
2.28
3.38
5.82
11.49
25.46
65.37
CacheBlend
2.11
2.15
3.86
4.12
7.48
19.73
59.78
RelayCaching
3.16
3.48
3.44
6.37
9.14
24.58
71.56
BaseShared
1.80
2.36
3.26
5.63
10.99
26.05
70.94
Appendix
Table 13: Single-stream TTFT in seconds and throughput in tokens per second under QKVO adaptation with r=4 on LLaMA-3.1-8B-Instruct. This configuration matches the LoRA parameter count of QV adaptation with r=8 , and TP denotes throughput.