Many-shot in-context learning (ICL) enables large language models (LLMs) to adapt to complex tasks by conditioning on thousands of demonstration examples, but this paradigm shifts the inference efficiency bottleneck to the key-value (KV) cache memory. Due to the linear scaling behavior of the KV cache, storing these intermediate tensors has become a paramount challenge for both online serving and on-device deployment. To address this issue, we propose a novel compression framework, termed MILO, that exploits the low-rank redundancy inherent in many-shot contexts. Specifically, MILO features a block-wise low-rank compression strategy that compresses the KV cache at the block granularity, where each block contains multiple many-shot examples. Furthermore, to handle the heterogeneous context density across different blocks, MILO dynamically allocates rank budgets based on the information entropy, preserving the fidelity of critical blocks while aggressively compressing redundant ones. Experimental results on Qwen2.5 models demonstrate that our method achieves up to 50% reduction in KV cache memory and 1.8x throughput improvement, with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
Figures & tables
Figure 1: Top-1% Effective rank of pre-RoPE key and value over different numbers of shots on the MathQA dataset.
Figure 2: Observation of the low-rank characteristics of the key and value cache in Qwen2.5 models across different blocks of examples on MATHQA datasets. Here we sample three blocks, where each block contains 64 examples. Effective rank is defined as the square of singular values. We calculate the mean and std across all attention layers.
Figure 3: Overview of MILO. During the encoding stage, the KV cache for many-shot contexts is split into N blocks, with each block containing m examples. MILO first applies SVD to compute the appropriate rank for each block based on entropy analysis, then performs subsequent low-rank compression with rank budgets (r1,...,rN) . During inference, the input query goes through the retriever to conduct KV cache selection, where the selected compressed KV cache is reconstructed in parallel to perform subsequent sparse attention.
Model
Method
BANKING77
CLINIC150
TREC
NLU
MATHQA
SVAMP
Avg.
Qwen2.5-3B
Full Cache
0.772
0.736
0.884
0.804
0.328
0.536
0.677
ASVD
0.722
0.694
0.812
0.723
0.241
0.441
0.605
Palu
0.740
0.712
0.856
0.768
0.264
0.484
0.637
MILO
0.744
0.720
0.860
0.780
0.276
0.500
0.647
Qwen2.5-7B
Full Cache
0.812
0.808
0.932
0.860
0.420
0.684
0.753
ASVD
0.785
0.742
0.837
0.815
0.311
0.594
0.680
Table 1: Accuracy comparison of MILO against baseline methods on various many-shot ICL tasks. All experiments are conducted using 512-shot examples under the same rank budget of 50%.
Figure 4: End-to-end system performance comparison of MILO against prior methods on the MathQA dataset under the same rank budget of 50% on a single NVIDIA L4 GPU.
Figure 5: Performance Analysis. Top Left (a): Impact of compression ratio (%) on model performance for MILO and baseline methods; Top Right (b): Impact of block size (number of examples per block) on latency and model accuracy for MILO.; Bottom Left (c): Performance breakdown for MILO customized Triton kernels; Bottom Right (d): Impact of rank allocation strategy on model accuracy for MILO. All experiments are conducted using Qwen2.5-7B.
Model
Method
KV Reduction
Accuracy
Qwen2.5-3B
Full Cache
1 ×
0.677
MILO
2 ×
0.647
MILO + INT8
4 ×
0.621
MILO + KIVI
4.7 ×
0.608
Qwen2.5-7B
Full Cache
1 ×
0.753
MILO
2 ×
0.725
Table 2: Accuracy and KV-cache memory reduction results when combining MILO with quantization.
The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. We propose AnchorKV, a compression scheme that shrinks the cache by 20× without discarding a single token. AnchorKV represents the cache using a small set of anchors stored exactly, expresses every other token through its most similar anchor, and refines only those whose approximation most affects the model's output. AnchorKV consistently preserves accuracy across models and datasets, retaining 99% of the full-cache score at the 70B scale, while keeping the entire context at a fraction of its cost.
Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2
Department of Computer Science, Technion - Israel Institute of Technology
As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.
Joao Monteiro, Louis Béthune, Anastasiia Filippova +3
Long-context LLM services now sustain prompts with hundreds of thousands to millions of tokens, making the key-value (KV) cache a first-order serving cost. Because the cache grows linearly with context length, it can exhaust GPU memory, force smaller batches, and reduce serving throughput. Prior KV cache compression techniques typically target only the sequence dimension or only the channel dimension, which leaves limited headroom as context windows scale. Compressing both dimensions promises higher memory reduction, but applying the two forms of compression directly leads to significant accuracy loss. This paper introduces MosaicKV, a dynamic two-D (dimensional) KV cache compression system for extremely long-context serving. MosaicKV uses dynamic two-D compression to address the accuracy challenge, exploiting the non-uniform importance distribution of elements within the KV cache. Instead of applying one compression pattern globally, MosaicKV identifies important elements for each KV vector and selects compression strategies at the granularity of KV cache segments. To address the performance challenge, where fine-grained sparsity and compression management overhead can offset the gains from compression, MosaicKV introduces compressed KV cache management. This mechanism uses underutilized GPU and CPU resources to maintain compressed KV caches and accelerate attention computation. Evaluation on an H800 GPU with multiple LLMs shows that MosaicKV delivers up to 16x attention speedup, 4.8x lower decode latency, and 7.3x higher throughput than the uncompressed baseline. At the same time, it reduces memory usage by 3x and incurs only 1.76% average accuracy loss on LongBench and RULER.
Sheng Qiang, Ruiwei Chen, Yinpeng Wu +5
Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University, Shanghai, China