cs.LGSep 27, 2026

Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving

Authors: Jiantong Jiang, Yue Yang, Peiyu Yang, Feng Liu

Organizations: University of Melbourne · Maincode

Abstract

Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready. Under memory pressure, this target-only contract can turn KV shortage into request stalls and preemptions. We present ElasticKV, a mixed-fidelity KV runtime built on the observation that target fidelity need not gate execution. ElasticKV introduces a compact intermediate KV state, making fidelity a runtime-managed execution property. To realize this state in a paged serving runtime, ElasticKV combines (i) a pair-structured layout that turns fidelity reduction into reusable GPU capacity, (ii) a dual-mode attention backend that directly consumes the compact state while preserving the native target-only path, and (iii) pressure-aware fidelity management that adapts KV fidelity to memory pressure. Our extensive evaluation across diverse workloads, model families and scales, and GPU platforms demonstrates the effectiveness and generality of ElasticKV. Under high concurrency, ElasticKV achieves 3.8-4.0×\times lower time-to-first-token (TTFT) and 9.1×\times lower P90 TTFT than vLLM while preserving generation quality.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving

    May 13, 2026Zedong Liu, Xinyang Ma, Dejun Luo +9Key-Value Cache CompressionLarge Language Model Serving

  2. MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

    Sep 7, 2026Michael Wang, Keith Li, Roozbeh BostandoostKey-Value Cache CompressionLLM Inference Optimization

  3. Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

    Jul 9, 2026Jiantong Jiang, Peiyu Yang, Rui Zhang +1Kv-Cache ManagementLarge Language Model Serving