LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it speeds up one deployment, slows down another, and changes nothing on a third, even where loading a token back is several times cheaper than recomputing it. The reason is that cached state must survive until it is used again. While one agent waits for its tool, the server processes the contexts of all other agents, so an agent's prefix is reused only if the host tier holds the reusable context of the whole agent pool, which we call the reuse working set. A smaller tier keeps writing state that is evicted before anyone reads it. We present EfficientAgent, which sizes and manages the host tier by this working set. A stack-distance model estimates the working set from agent histories to size the host tier; its predictions, made before the experiments, located the capacity at which offloading starts to pay. When the tier is too small, a runtime policy stops writing large refills of evicted context and keeps extending prefixes that are still cached; when the tier is large enough, it writes everything. On SWE-bench Verified coding agents, a host tier sized to the estimated working set cuts recomputed prompt tokens by 93% and end-to-end time by 39%. With a small fixed tier, the policy cuts recomputation by 35%; with a large tier, it avoids the 4.3-fold increase caused by always filtering writes. Across three GPU types and two models, offloading pays off when the GPU has little compute per byte of host bandwidth and the host tier holds the working set. Code is available at https://github.com/KunmingSHAO/efficientagent_release.
Figures & tables
GPU
BF16 TFLOPS
Host link (GB/s)
Memory (GB)
FLOP/byte
Offload/recompute
RTX 3090
71
PCIe Gen4, 32
24
2.2K
0.91 ; dense 0.92
H20
148
PCIe Gen5, 64
96
2.3K
0.60 (tier ≥ working set)
H800
989.5
PCIe Gen5, 64
80
15.5K
1.08–1.87; dense 1.23
Table 1: Offload pays on low-ratio GPUs once host hits survive and slows every H800 run. Peak dense BF16 TFLOPS, host link per direction, and memory per GPU; offload/recompute for the MoE and the dense model. Bold: offload faster than recomputation.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Model / precision
Qwen3-Coder-30B-A3B-Instruct / BF16
Maximum model context
262,144 tokens
Running-request cap
16 ( max_num_seqs )
Scheduler token budget
8,192; chunked prefill enabled
GPU prefix caching
Enabled
Active pool A
16; 8 in the concurrency comparison
Appendix
Table 2: Serving configuration. Upper part: the H20 stack; the active pool limits the tasks in progress, and the running-request cap limits the requests the engine executes at once. Lower part: memory configuration per deployment of Qwen3-Coder-30B-A3B-Instruct and of the dense Qwen2.5-Coder-32B-Instruct; GPU KV capacity is in tokens, each GPU holding its shard of every token.
Configuration
A
CPU GiB
Time (min)
Prefill M
Retrieved M
Stored M
Preempt.
Recompute
16
0
211.69
89.66
0
0
2,190
Recompute, repeat
16
0
205.98
89.82
0
0
2,203
Offload
16
5
208.91
76.74
12.57
81.69
1,808
Offload
16
10
161.06
23.53
65.20
22.42
437
Offload
16
20
128.05
5.28
82.77
4.06
46
Offload
16
40
124.01
5.27
82.43
4.06
52
Appendix
Table 3: Complete H20 replay measurements. Every completed row processes the same 4,427 calls. Each row is one run; repeats are retained individually ( ∗ : see Baselines and policy variants). A : active pool; host (CPU) budgets are per rank. Retrieved and stored token volumes are summed over TP workers and divided by eight. Preempt.: engine preemptions during the replay. Bold: capacity-conditioned admission.
Host GiB
Predicted prefill
Measured prefill
Predicted restore
Measured restore
10
27.64–31.20
23.53
55.28–58.88
65.20
20
5.19–5.78
5.28
80.84–81.36
82.77
80
5.19–5.31
5.29
81.31–81.36
82.89
Appendix
Table 4: Prefix-survival predictions and subsequent measurements. Prefill and restored volumes are millions of tokens; restored volume is per rank. Prediction ranges use the two reference interleavings.
Host GiB
Predicted prefill
Measured prefill
Predicted restore
Measured restore
0
41.87–60.80
41.48
0
0
3
29.59–52.01
28.99
7.51–12.28
13.01
4.5
18.84–31.73
17.11
22.25–29.36
13.22
Appendix
Table 5: Predictions and measurements for an active pool of A=8 . Units match Table 4 .
Qwen3-Coder-30B-A3B-Instruct
Qwen2.5-Coder-32B-Instruct
TP ranks
Heads/rank
KiB/rank
Aggregate KiB
Heads/rank
KiB/rank
Aggregate KiB
1
4
96
96
8
256
256
2
2
48
96
4
128
256
4
1
24
96
2
64
256
8
1
24
192
1
32
256
Appendix
Table 6: BF16 KV footprint of the two evaluated models.
Parameter
Value
Cache occupancy threshold θ
0.95
Task activity and eviction window
60 seconds
Prompt-length window
Last 256 first prefix lookups
Telemetry report interval
1 second
Scheduler report-read interval
0.5 seconds
Maximum age of a fresh report
5 seconds
Appendix
Table 7: Write-admission parameters, fixed before the admission runs.
Hardware
Model
Recompute
Offload
Ratio
Workflow wall-clock
8× RTX 3090
MoE
407 min
369 min
0.91
8× RTX 3090
Dense
443 min
406 min
0.92
8× H20
MoE
20.58 h
20.80 h
1.01
Engine window
2× H800
MoE
62 min
73 min
1.18
Appendix
Table 8: Offload shortens the RTX 3090 runs and lengthens every H800 run for both models. Both configurations enable GPU prefix caching. Rows are grouped by timing scope; each ratio (offload/recompute) compares the two within one deployment and scope. MoE: Qwen3-Coder-30B-A3B-Instruct; dense: Qwen2.5-Coder-32B-Instruct. † Complete SWE-bench Verified benchmark. Bold: offload faster than recomputation.
Hardware
Without prefix caching
With prefix caching
Speedup
8× RTX 3090
13 h 35 min
1 h 13 min
11.1 ×
2× H800
1 h 58 min
54 min
2.19 ×
Appendix
Table 9: Prefix caching shortens coding-agent runs by up to 11.1 × . Qwen3-Coder-30B-A3B-Instruct; workflow wall-clock.
Recompute
Offload
Hardware
GPU KV usage
Total
Total
GPU
External
8× RTX 3090
< 80%
62.9
79.4
62.7
16.6
80–90%
34.7
43.7
16.9
26.7
90–95%
34.4
44.0
17.1
27.2
95–100%
33.8
44.3
17.3
25.9
2× H800
< 80%
65.7
98.4
97.5
0.2
Appendix
Table 10: Prefix-hit medians by GPU KV usage. The Qwen3-Coder-30B-A3B-Instruct RTX 3090 and two-H800 subset pairs of Table 8 , both at GPU memory fraction 0.65. Entries are medians (%) within each usage band. Without offload, all prefix hits are GPU hits. With offload, the GPU, external, and total rates are separately computed medians; the GPU and external medians need not sum to the total. Bold: net hit gain with offload at high GPU KV usage.
Table 13: Condensers that rewrite early history lower the prefix-hit rate. Qwen3-Coder-30B-A3B-Instruct on 8× RTX 3090 with GPU prefix caching; run means.
Agentic serving can consume orders of magnitude more tokens than chatbot workloads, stressing both KV-cache capacity and decode-time bandwidth. Most KV-eviction methods score cached keys against representative queries drawn from the most recent tokens, assuming future attention resembles recent attention. We show that agentic generation violates this assumption: future queries form a mixture over think, act, tool, and others phases, and principal-angle analysis shows these components occupy measurably different query subspaces, so recency representatives systematically undervalue keys that upcoming phases will need. We propose AGENTKV, which maintains a small query buffer per phase and scores cached keys against their union. We further implement AGENTKV in a persistent multi-turn serving path that carries compressed KV state across turns and compacts retained KV pages online. Across two models, six task domains, and three KV budgets each, AGENTKV improves task score by 5.5 points on average over R-KV and 5.3 over Tri-attention. Relative to upstream full-KV SGLang, AGENTKV improves output-token throughput by up to 1.80x. Code: https://github.com/LiuTaowen-Tony/agentkv.
Taowen Tony Liu, Jeffrey T. H. Wong, Can Xiao +3
Imperial College London, London, United Kingdom · Columbia University, New York, NY, USA
LLM agents today communicate via text, which incurs considerable latency and information loss due to the need to autoregressively decode the sharer model's state and encode at the receiver model. Recent work such as Cache-to-Cache (C2C; Fu et al., 2026) seeks to exchange KV caches by learning adapters that translate sharer KV matrices to the receiver model. However, the adapters are large and expensive to train, and translate individual tokens, which requires the target context to be identical. This is unsuitable for agent communication, where the LLMs have differing context. We introduce Latent Cache Flow (LCF). To address efficiency, we observe that keys and values can be jointly translated and compressed, reducing the adapter to about 4% of C2C's size. To address differing context, we design the adapter to transmit a summary of new information that the target model does not have. Our early experiments show that a pruned 13 MB LCF adapter can be more accurate than C2C at 956 MB in shared-context settings; for different contexts, LCF improves F1 by 7.5% and Exact Match by 23% while 8.5 times faster than text-based communication.
LLM serving caches prompt KV state, yet most front ends still re-tokenize the full request on every call. Coding agents pay most: sessions repeatedly submit a long transcript after a small append, which can shift token boundaries near the end of the prior sequence. Across 153,951 calls the median append is ~1.4K characters; only 1.0-3.6% of calls start or rebuild a session, yet those carrymulti-million-character contexts. Fleet prompt-cache hit rate is 94.1%, and as it approaches 0.99, tokenization grows from 10% to 64% of time to first token (TTFT) in component measurements. TokTier is a stateful CPU+GPU tokenization service for this two-mode workload, under one contract: emitted token IDs are always identical to full reference tokenization. For session continuations it re-tokenizes a small window around the append and splices only when a per-request check finds a stable pre-tokenization boundary; failed checks widen the window or fall back to full reference tokenization. For calls without a reusable prefix it runs exact GPT-family regex pre-tokenization and BPE on a GPU. A sampled shadow verifier re-checks live traffic. Across 17 production tokenizer families, differential campaigns cover 1.5x10^10 split checks, a 12.4 TB real-text corpus, and 93,000+ replayed agent steps, with zero divergence. Incremental repair takes 0.5-1.1 ms from 100K to 3M characters, up to 437x faster than HF tokenization and 2.1x faster at 1M characters than the strongest cache-based baseline (Gigatoken) fully prewarmed. GPU tokenization encodes a 1M-character request in 0.87 ms, up to 491x below HF and 23.4x below the fastest published CPU method on the same protocol. With vLLM, median TTFT drops 16-34% and P99 TTFT 23% under recorded bursts. Under a 50 ms P99 objective, a four-core repair pool plus one GPU sustains 1,821 requests/s, where a 16-core stateless front end saturates at 40 requests/s.