Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving
Organizations: University of Melbourne · Maincode
Abstract
Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready. Under memory pressure, this target-only contract can turn KV shortage into request stalls and preemptions. We present ElasticKV, a mixed-fidelity KV runtime built on the observation that target fidelity need not gate execution. ElasticKV introduces a compact intermediate KV state, making fidelity a runtime-managed execution property. To realize this state in a paged serving runtime, ElasticKV combines (i) a pair-structured layout that turns fidelity reduction into reusable GPU capacity, (ii) a dual-mode attention backend that directly consumes the compact state while preserving the native target-only path, and (iii) pressure-aware fidelity management that adapts KV fidelity to memory pressure. Our extensive evaluation across diverse workloads, model families and scales, and GPU platforms demonstrates the effectiveness and generality of ElasticKV. Under high concurrency, ElasticKV achieves 3.8-4.0 lower time-to-first-token (TTFT) and 9.1 lower P90 TTFT than vLLM while preserving generation quality.
Figures & tables
| Setting | Llama-3.1-8B | Qwen3-8B |
|---|---|---|
| Static FP16 | 49.89 | 49.59 |
| Static FP8 | 48.72 | 49.07 |
| ElasticKV (5% cap) | 50.06 | 49.59 |
| ElasticKV (10% cap) | 50.06 | 49.52 |
| ElasticKV (20% cap) | 50.04 | 49.42 |
| All-BASE | 48.45 | 47.87 |
| Activity | Cost (ms/step) |
|---|---|
| Online control | 0.32 |
| Sync. demotion | 2.82 |
| Elastic metadata | 1.23 |
| Variant | P90 TTFT | P90 TPOT | BASE |
|---|---|---|---|
| ElasticKV | 2.8s | 40.3ms | 14.7% |
| Watermark demotion | 3.4s | 46.5ms | 17.2% |
| Eager promotion | 4.4s | 52.0ms | 15.7% |
| No promotion | 2.6s | 40.5ms | 28.8% |
| General elastic path | 3.1s | 49.5ms | 14.7% |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Reference settings | ElasticKV exposure cap | |||||||||
| Task | Metric | FP16 | FP8 | All-BASE | ||||||
| Score | Exp. | Score | Exp. | Score | Exp. | |||||
| Single-Document QA | ||||||||||
| NarrativeQA | QA F1 | 27.85 | 28.34 | 27.61 | 28.87 | 5.3% | 28.95 | 6.8% | 28.95 | 6.8% |
| Qasper | QA F1 | 44.76 | 43.81 | 44.88 | 44.83 | 5.2% | 44.69 | 8.6% | 44.61 | 20.5% |
| MultiFieldQA-en | QA F1 | 55.62 | 53.38 | 55.48 | 56.01 | 3.7% | 56.09 | 10.5% | 56.08 | 20.5% |
| Reference settings | ElasticKV exposure cap | |||||||||
| Task | Metric | FP16 | FP8 | All-BASE | ||||||
| Score | Exp. | Score | Exp. | Score | Exp. | |||||
| Single-Document QA | ||||||||||
| NarrativeQA | QA F1 | 26.32 | 24.39 | 24.73 | 26.88 | 4.8% | 26.64 | 6.7% | 26.64 | 6.7% |
| Qasper | QA F1 | 47.70 | 45.87 | 45.08 | 47.36 | 5.5% | 47.03 | 10.3% | 47.4 | 18.0% |
| MultiFieldQA-en | QA F1 | 53.34 | 54.34 | 51.33 | 52.83 | 3.9% | 52.93 | 10.5% | 52.56 | 19.6% |
| Model | System | Preemptions | BASE exposure |
|---|---|---|---|
| Llama-3.1-8B-Instruct | Full-GPU | 1297 | – |
| Exact-offload | 1392 | – | |
| ElasticKV | 0 | 2.6% | |
| Qwen3-8B | Full-GPU | 1574 | – |
| Exact-offload | 1594 | – | |
| ElasticKV | 0 | 3.6% |
| Window | Region | Trace interval | Requests | Input tokens | Output tokens |
|---|---|---|---|---|---|
| A | Early | [450,750) | 920 | 11.82M | 0.324M |
| B | Middle | [1650,1950) | 1027 | 11.54M | 0.357M |
| C | Late | [2850,3150) | 1122 | 13.37M | 0.358M |
| Window | Region | Trace interval (s) | Requests | Input tokens | Output tokens |
|---|---|---|---|---|---|
| A | Early | [300,900) | 1810 | 24.34M | 0.627M |
| B | Middle | [1500,2100) | 2093 | 23.45M | 0.731M |
| C | Late | [2700,3300) | 2207 | 25.48M | 0.720M |
| Window | System | Eff@300 | Mean TTFT | P90 TTFT | Preempt. | Drain | BASE exp. (%) |
|---|---|---|---|---|---|---|---|
| A | Full-GPU | 27.79% | 272.49s | 924.63s | 649 | 1454s | – |
| Exact-offload | 25.47% | 84.42s | 251.64s | 450 | 665s | – | |
| ElasticKV | 38.12% | 40.16s | 118.63s | 0 | 299s | 4.84 | |
| B | Full-GPU | 36.69% | 96.79s | 356.03s | 484 | 477s | – |
| Exact-offload | 35.88% | 24.89s | 76.60s | 116 | 403s | – | |
| ElasticKV | 55.90% | 12.89s | 39.73s | 0 | 109s | 0.01 |
| Model | Setting | Executable KV states | Backend | P90 TTFT (s) | Out. tok/s | Preempt. |
|---|---|---|---|---|---|---|
| Llama-3.1-8B | Static FP8 | TARGET FP8 | FlashInfer | 1.47 | 1003.12 | 0 |
| ElasticKV | TARGET 16 + BASE 8 | Elastic Triton | 2.87 | 914.25 | 0 | |
| Qwen3-8B | Static FP8 | TARGET FP8 | FlashInfer | 1.50 | 949.04 | 0 |
| ElasticKV | TARGET 16 + BASE 8 | Elastic Triton | 1.98 | 854.92 | 0 |
| Activity | Trigger (%) | Cost (ms/step) |
|---|---|---|
| Online control | ||
| Shortage probe | 78.67 | 0.072 |
| Demotion planning | 11.06 | 0.175 |
| Promotion planning | 100.00 | 0.023 |
| Restore-source prep. | 100.00 | 0.021 |
| Async promotion issue | 13.63 | 0.029 |
| Statistic | Value |
|---|---|
| Scheduling steps | 16586 |
| BASE exposure | 14.71% |
| Successful demotion batches | 1835 |
| Demoted pairs | 29427 |
| Promotion issues | 2260 |
| Promoted pairs | 25278 |
| Metric | Profile off | Profile on |
|---|---|---|
| Duration (s) | 673.01 | 673.08 |
| Mean TPOT (ms) | 39.39 | 39.38 |
| Output throughput (tok/s) | 912.92 | 912.82 |
| Native preemptions | 0 | 0 |