Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.
Figures & tables
Method
Accuracy Preservation
Repeated Prefill Reduction
Existing Adapter Support
Selective Recomputation
✗
✓
✓
BaseShared
✓
✗
✓
BaseLRShared
✓
✓
✗
PReCache (Ours)
✓
✓
✓
Table 1: Conceptual comparison of KV cache sharing strategies for multi-LoRA agents across three criteria. ✓ and ✗ denote support and no support, respectively.
Figure 1: Overview of KV cache construction in PReCache across consecutive agent turns. (A1–A3) PreLRShared uses the hidden states generated when each context segment is first processed to construct the LR caches for all agents, eliminating the separate full-length backbone processing over the accumulated context LP required by BaseShared. (B1–B5) ReBaseShared reconstructs the shared base cache from adapter-free hidden states to reduce dependency on the previous agent’s adapted representation. ReBaseShared DB (B2) schedules the adapted and adapter-free paths together during both prefill and decoding, whereas ReBaseShared LP (B3–B4) performs adapter-free reconstruction as a contiguous prefill after the current agent’s turn.
Figure 2: Effect of reconstructing the shared base cache from adapter-free hidden states in Ministral-8B on a HotpotQA trajectory. (a) compares the previous agent’s and neutral base caches against the current agent’s base cache. (b) reports the layer-wise relative errors in the hidden states and base caches of PreLRShared and ReBaseShared. (c) reports the last-layer relative error across agent transitions and includes FullShared as a reference for entire KV cache reuse.
HotpotQA
ScienceQA
Model
Method
Easy
Medium
Hard
Avg.
1–4
5–8
9–12
Avg.
LLaMA-3.1-8B
NonShared
40.55
41.05
30.95
37.52 0.00
70.33
59.67
77.58
69.19 0.00
FullShared
38.05
37.90
26.15
34.03 − 3.48
67.71
56.42
72.75
65.62 − 3.57
DroidSpeak
39.80
38.30
27.60
35.23 − 2.28
68.17
59.04
74.92
67.37 − 1.82
CacheBlend
39.15
38.55
27.10
34.93 − 2.59
67.90
58.38
73.71
66.66 − 2.53
RelayCaching
40.15
39.55
29.20
36.30 − 1.22
68.79
59.25
75.83
67.96 − 1.23
Table 2: Mean benchmark accuracy (%) of NonShared and KV cache sharing methods on HotpotQA and ScienceQA. The value beside each average denotes its percentage-point difference from NonShared. For each backbone and benchmark, the higher and lower halves of the methods ranked by average accuracy are highlighted in green and red, respectively.
Figure 3: Single-stream TTFT ( ↓ ) and per-request throughput ( ↑ ) over trajectory length for both backbones on a single A6000 GPU. ReBaseShared uses lazy prefill scheduling.
Figure 4: Serving TTFT percentiles ( ↓ ) and per-request throughput ( ↑ ) over request rate on vLLM with a fixed 17.3 k-token trajectory on a single A100 GPU. ReBaseShared uses double batching.
Figure 5: Efficiency comparison of LP and DB scheduling for ReBaseShared. The left panels report single-stream TTFT ( ↓ ) and per-request throughput ( ↑ ) over trajectory length. The right panels report p50 TTFT ( ↓ ) and per-request throughput ( ↑ ) over request rate with a fixed 17.3 k-token trajectory under concurrent serving.
Figure 6: Peak GPU memory usage across trajectory lengths for LLaMA-3.1-8B.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Layer-wise geometry of neutral base cache reconstruction in ReBaseShared. Each quantity is computed per token and then averaged across tokens, prompts, and agent transitions. Solid and dashed lines represent base cache and hidden state measurements, respectively. At the token level, ρ>2cosθ , or equivalently Γerr>1 , indicates lower relative error for the neutral state than for the state constructed from the previous agent’s hidden states.
Method
TTFT (s)
Model Latency (s)
E2E Latency (s)
Trajectory Length (tokens)
Average
Maximum
NonShared
1.41
7.23
16.21
1268
6857
FullShared
0.90
9.20
20.64
1757
7078
DroidSpeak
1.29
7.10
16.11
1340
6317
CacheBlend
1.35
6.98
16.64
1569
7103
RelayCaching
1.31
8.00
17.58
1455
6982
Appendix
Table 3: Average benchmark latency and trajectory length for LLaMA-3.1-8B on HotpotQA. Latencies are reported in seconds, with average and maximum trajectory lengths reported in tokens.
Figure 8: Accuracy-latency trade-off for LLaMA-3.1-8B on HotpotQA. The panels compare benchmark accuracy with TTFT, model latency, and end-to-end latency, respectively. Higher accuracy and lower latency are preferred.
Benchmark
NonShared
FullShared
DroidSpeak
CacheBlend
RelayCaching
BaseShared
PreLRShared
ReBaseShared
HotpotQA
0.21
0.44
0.28
0.34
0.31
0.25
0.30
0.32
ScienceQA
0.28
0.60
0.34
0.35
0.39
0.35
0.39
0.37
Appendix
Table 4: Standard deviation of benchmark accuracy across 20 complete evaluation runs, reported in percentage points.
Model
Metric
Method
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
LLaMA-3.1-8B
TTFT (s)
NonShared
1.93
2.53
3.72
6.77
16.36
33.72
73.18
FullShared
1.14
1.27
1.63
2.39
4.22
9.05
23.15
DroidSpeak
1.62
2.15
3.21
5.55
11.12
25.23
67.82
CacheBlend
1.67
2.15
2.31
3.72
8.61
19.05
56.38
RelayCaching
1.87
2.35
2.44
4.67
9.89
22.07
71.19
BaseShared
1.61
2.13
3.08
5.25
10.54
23.80
68.19
Appendix
Table 5: Single-stream TTFT in seconds and throughput in tokens per second over trajectory length on a single A6000 GPU. ReBaseShared uses LP scheduling, and TP denotes throughput.
Metric
Method
QPS
0.5
1
2
4
8
16
TP (tok/s)
NonShared
1573
1282
596
70
52
48
FullShared
2922
2910
2818
1809
205
164
DroidSpeak
1592
1416
686
78
60
53
CacheBlend
1939
1847
1391
140
104
83
RelayCaching
1827
1678
835
93
68
59
Appendix
Table 6: Serving TTFT percentiles in seconds and per-request throughput in tokens per second over request rate with a fixed 17.3 k-token trajectory on a single A100 GPU. ReBaseShared uses DB scheduling, and TP denotes throughput.
Metric
Schedule
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
TTFT (s)
LP
1.23
1.36
1.94
3.19
6.06
12.53
27.24
DB
1.24
1.58
2.33
3.94
7.72
17.95
49.29
TP (tok/s)
LP
166
245
404
635
900
1087
1068
DB
126
189
299
486
723
907
889
Appendix
Table 7: LP and DB scheduling for ReBaseShared under single-stream inference on LLaMA-3.1-8B. TP denotes throughput.
Metric
Schedule
QPS
0.5
1
2
4
8
16
p50 TTFT (s)
LP
0.06
0.06
6.97
18.26
22.67
24.79
DB
0.07
0.07
0.21
1.29
4.84
6.83
TP (tok/s)
LP
1670
1371
947
69
47
38
DB
2036
1972
1408
182
103
85
Appendix
Table 8: LP and DB scheduling for ReBaseShared under concurrent serving with a fixed 17.3 k-token trajectory. TP denotes throughput.
Method
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
NonShared
15.69
16.07
16.83
18.33
21.35
27.39
39.47
FullShared
15.29
15.47
15.85
16.60
18.10
21.11
27.13
DroidSpeak
15.52
15.86
16.48
17.75
20.29
25.37
35.50
CacheBlend
15.49
15.77
16.36
17.52
19.86
24.52
33.86
RelayCaching
15.46
15.73
16.28
17.38
19.59
24.01
32.84
BaseShared
15.30
15.49
15.88
16.66
18.22
21.34
27.58
Appendix
Table 9: Peak GPU memory usage over trajectory length on LLaMA-3.1-8B.
Metric
Method
N=4
N=8
TTFT (s)
NonShared
68.36
90.05
FullShared
23.98
22.73
DroidSpeak
66.57
88.89
CacheBlend
48.71
72.13
RelayCaching
52.99
80.96
BaseShared
58.09
87.60
Appendix
Table 10: Efficiency and peak GPU memory usage for different numbers of agents on LLaMA-3.1-8B-Instruct using a single A6000 GPU. Each turn adds 8 k tokens over a total of eight turns.
Method
r=4
r=8
r=16
r=32
NonShared
34.72
36.72
36.70
36.80
RelayCaching
32.95
34.63
34.47
34.20
PreLRShared
33.51
34.77
34.67
34.53
ReBaseShared
34.41
36.08
35.98
35.81
Appendix
Table 11: HotpotQA accuracy (%) across LoRA ranks on Ministral-8B-Instruct.
Method
r=8
r=16
r=32
r=64
NonShared
768
762
756
750
FullShared
1693
1696
1694
1690
DroidSpeak
937
927
921
913
CacheBlend
1042
1049
1042
1033
RelayCaching
900
902
891
878
BaseShared
976
967
958
953
Appendix
Table 12: Single-stream throughput in tokens per second across LoRA ranks at a fixed 33.7 k-token trajectory on LLaMA-3.1-8B-Instruct.
Metric
Method
1.9k
3.0k
5.0k
9.1k
17.3k
33.7k
66.4k
TTFT (s)
NonShared
1.81
4.16
4.39
6.45
12.91
27.01
68.73
FullShared
1.45
1.45
2.00
2.71
5.23
10.18
24.62
DroidSpeak
1.76
2.28
3.38
5.82
11.49
25.46
65.37
CacheBlend
2.11
2.15
3.86
4.12
7.48
19.73
59.78
RelayCaching
3.16
3.48
3.44
6.37
9.14
24.58
71.56
BaseShared
1.80
2.36
3.26
5.63
10.99
26.05
70.94
Appendix
Table 13: Single-stream TTFT in seconds and throughput in tokens per second under QKVO adaptation with r=4 on LLaMA-3.1-8B-Instruct. This configuration matches the LoRA parameter count of QV adaptation with r=8 , and TP denotes throughput.
Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and construct a separate cache with high computation and memory overhead. Selective recomputation reduces this redundancy but still retains substantial model execution, while existing delta correction methods either support only recurring context relations or maintain memory-intensive online correction states for dynamically changing context. For first seen shared context, these methods also construct a reference cache outside the agent workflow, and an approximate correction at the first agent affects the outputs passed to subsequent agents. We present KVCMAS, an online KV cache correction framework that represents cross-agent cache deviations using compact low-rank states and seamlessly chains corrections along the agent workflow without an additional reference prefill. This design supports dynamically changing shared context while preserving an exact first-agent cache. Across multiple language and vision-language workloads, KVCMAS matches or improves the accuracy of prior KV cache sharing methods while achieving the lowest TTFT under highly concurrent serving. Under controlled serving traces, it provides a 2.0x TTFT speedup over inference without KV cache sharing and reduces peak GPU memory by up to 3.7x relative to a prior KV cache correction method. These results establish KVCMAS as an accurate and scalable KV cache sharing approach for prompt-specialized multi-agent serving.
Hyesung Jeon, Hyeongju Ha, Seoyoung Lee +2
Department of Electrical and Computer Engineering Seoul National University
We present PolyKV, a system in which multiple concurrent inference agents share a single, asymmetrically compressed KV cache pool. Rather than allocating a separate KV cache per agent -- the standard paradigm -- PolyKV writes a compressed cache once and injects it into N independent agent contexts via HuggingFace DynamicCache objects. Compression is asymmetric: Keys are quantized at int8 (q8_0) to preserve softmax stability, while Values are compressed using TurboQuant MSE -- a Fast Walsh-Hadamard Transform (FWHT) rotation followed by 3-bit Lloyd-Max quantization with centroids tuned to N(0,1). We evaluate across two model scales (SmolLM2-1.7B-Instruct and Llama-3-8B-Instruct), three context lengths (600-7,194 tokens), and up to 15 concurrent agents. PolyKV achieves a stable 2.91x compression ratio across all configurations. On Llama-3-8B with 15 agents sharing a 4K-token context, PolyKV reduces KV cache memory from 19.8 GB to 0.45 GB -- a 97.7% reduction -- while maintaining only +0.57% perplexity degradation and a mean BERTScore F1 of 0.928. PPL delta does not grow with agent count and improves as context length increases, inverting to -0.26% at 1,851 coherent tokens. To our knowledge, no prior work combines a single shared, lossy-compressed KV pool with multi-reader concurrent agent access.
A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefills that shared context once per specialist. We study a narrow, practical question: for already-trained standard LoRA adapters -- not adapters retrained for cache compatibility -- how much task quality is preserved if the backbone's prefill KV cache is computed once and reused across specialists, and what does that buy in serving cost? On a Qwen3-1.7B backbone with two adapters (extractive QA on HotpotQA, arithmetic reasoning on GSM8K), we sweep the boundary at which the specialist takes over from the reused base cache and measure paired quality differences and serving cost. Full-prefix reuse had the lowest prefill cost and a small quality difference on held-out GSM8K (Delta = -4.6 EM at a 160-token budget; -3.0 at 320 tokens; -0.8 under a second training seed -- all favoring native, only the first excluding zero, and the magnitude not consistent). Partial recomputation provided no demonstrated advantage. Neither quality equivalence nor a general boundary-selection rule is established. We also report a closed-form ridge KV translator that did not beat direct reuse, and specialist-dependence contrasts whose intervals all include zero. The measured serving benefit is warm-cache time-to-first-token, which grows with context (~16x at 8K); two-branch peak memory was only 12% lower and, on inspection, the prefix was never physically shared across branches -- this implementation reuses KV values but copies their storage, so shared-cache memory savings are not achieved.