cs.LGJul 23, 2026

Error Certificates for KV-Cache Eviction via Randomized Design

Authors: Peng Xie

Organizations: Technical University of Munich

Abstract

Deterministic KV-cache eviction keeps the top-kk tokens under an importance score and deletes the rest. We prove that this design cannot know what it destroyed: evicted values can be altered so that everything the serving system retains is unchanged while the true attention-output error grows arbitrarily, so no serving-time estimator of that error is consistent. Randomized eviction restores identifiability. With a Poisson-sampled tail at known inclusion probabilities, one logit offset performs the Hájek correction inside the softmax, and a survey-sampling variance estimator over the retained set becomes a per-step error certificate with 0.97 empirical coverage at no accuracy cost. On real workloads, seven pre-registered claims locate the certificate's value precisely. Prediction goes to output confidence: question-aware eviction at 25--50% budgets is nearly free, output log-probability predicts failure better than any cache-side signal, and certificate-gated budget escalation adds nothing. Attribution stays with the certificate: it separates cache-induced from inherent failures (AUC 0.65--0.75, against 0.47--0.54 for output confidence) and schedules recomputation better than random or confidence gating. Randomization buys attribution, not prediction.

Explore similar work

Jul 30, 2026cs.AR

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization

KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the request it is serving right now. We give it a provably sound runtime meter -- a "DTrace for KV quantization": a per-(layer, head, step) upper bound on the total variation between exact and compressed attention. The meter has two tiers: a deterministic band-norm-witness bound, sound for any cache-preserving black-box quantizer and for any query (adaptive-safe, worst-case Cauchy--Schwarz plus RoPE band-unitarity), and a tighter probabilistic certificate for a controlled subtractively-dithered INT8 quantizer under an explicit request-level failure budget (stated for non-adaptive queries; core theorems machine-checked in Lean 4). Three results. Observability: the meter enters SGLang through an env-guarded patch, and any scheme registered as one tensor function is measured in live serving. Repair: meter-driven gating -- risk-ranked where the witness is saturated, certified where it is informative -- empirically restores the quality floor at benchmark scale, e.g. raw-cast fp8 from 22.8 back to 79.7 on hard RULER tasks with the difference from uncompressed bounded at [+0.0,+0.8][+0.0,+0.8] by a paired test. Analysis: aggressive schemes survive on cross-layer error cancellation, not per-step fidelity -- in a 28-layer sweep, no single layer's pollution alone loses anything (0/28) -- and the certified int8 cache serves 1.88×1.88\times more KV tokens at the same memory in SGLang. All artifacts, guards, and the Lean development are released at https://github.com/metask-ai/witcert-kv-certificates; every number regenerates from the shipped artifacts by one command.
Fanzhe Wei, Li Liu, Ziyang Wang +1
Aug 26, 2026cs.LG

PAGE: Partition-Aware Gated KV-Cache Eviction

KV-cache eviction can do more than compress. In long-context LLMs, keeping only some cached tokens sometimes matches or exceeds full-cache accuracy, because many redundant prefill tokens otherwise dilute attention away from the tokens that carry the answer. This benefit is not uniform, and evicting the wrong tokens can drop accuracy to zero on tasks that require precise retrieval, so the useful question is not only which tokens to keep but also whether to evict this input at all. We show that one label-free number computed from the prefill attention, the drop between early and late layers in how much attention heads agree on which tokens to read, predicts per input, before any decoding, which of the two cases an input falls under. We build this into PAGE (Partition-Aware Gated Eviction), a wrapper that runs any SnapKV-style evictor when the drop is large and keeps the full cache when it is small, with no training, labels, or fine-tuning. PAGE is a safety mechanism rather than a compressor, so we measure it by the failures it prevents. It cuts the harm rate on capacity-bound inputs from 0.75 to 0.026, and on multi-key retrieval with Mistral-7B plain SnapKV falls from 99% to 0% as the budget shrinks, while PAGE holds it at 89%. Elsewhere, it passes the base evictor through unchanged, which is the intended behaviour and is what we observe in 8 of 16 cells. Code is available at https://anonymous.4open.science/r/PAGE-018239.
Pankaj Kumar, Subhankar Mishra
Jun 25, 2026cs.LG

Epiphany-Aware KV Cache Eviction Without the Attention Matrix

As reasoning models emit chains of thought tens of thousands of tokens long, KV cache increasingly becomes a deployment bottleneck. Existing cache eviction methods rank tokens by attention weight, which is a noisy importance proxy in long reasoning traces, and prohibits the use of fused kernels in production inference by forcing the model to materialize the attention matrix. In this work, we instead score tokens with a metric we term the epiphany score: the change in the model's internal representation, read directly from the forward pass with no attention matrix and negligible extra state. Our resulting cache eviction method, EpiKV, requires no training, classifier, or custom kernel, and can be used directly in FlashAttention inference stacks unchanged -- scaling to a 16x longer feasible context than attention-based scoring. upper-mid layers negatively) and remove a positional trend with a causal rolling z-score. At a 4096-token cache EpiKV reaches 72% on MATH-500, matching the strongest attention-based baseline (ThinKV 71%, H2O 67%); a lag-normalized KV variant reaches 37% on AIME-2024 at 8192 tokens against the best of them (33%), at up to 2.8x the speed.
Steven Kolawole, Virginia Smith