SparseEngine: Sparse-First Inference Engine
Abstract
Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specific layouts or workflows. We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure. SparseEngine supports 15 methods across four categories and enables cross-request state management through Chain Cache, which resumes KV-eviction methods from retained history, and controllable Prefix-Cache Pruning, which removes KV from selected history regions while preserving logical-prefix matching. While maintaining method quality, SparseEngine delivers over 10x higher throughput with KV eviction, over 2.5x faster decoding at matched concurrency than vLLM, and over 2x end-to-end speedup on agent benchmarks. The code is available at https://github.com/CURRENTF/SparseEngine.
Figures & tables
| LongBenchV1 | LongBenchV2 | |||||||||||||
| Method | S-Doc | M-Doc | Summ. | F-Shot | Synth. | Code | Avg. | S-Doc | M-Doc | Hist. | Learn. | Struct. | Code | Avg. |
| GLM-4.7-Flash | 41.6 | 47.1 | 27.8 | 71.0 | 52.5 | 65.7 | 51.0 | 34.3 | 28.8 | 20.5 | 29.6 | 21.2 | 46.0 | 31.4 |
| SnapKV | 41.0 | 47.2 | 26.9 | 69.8 | 52.5 | 65.2 | 50.4 | 36.0 | 31.2 | 20.5 | 32.1 | 30.3 | 42.0 | 33.2 |
| OmniKV | 41.7 | 47.1 | 27.8 | 70.6 | 52.5 | 65.1 | 50.8 | 35.4 | 28.0 | 23.1 | 33.3 | 42.4 | 42.0 | 33.4 |
| Quest | 41.4 | 47.2 | 27.7 | 71.2 | 52.5 | 65.4 | 50.9 | 34.9 | 29.6 | 23.1 | 40.7 | 27.3 | 42.0 | 33.8 |
| StreamingLLM | 34.7 | 41.7 | 26.4 | 69.2 | 53.0 | 64.1 | 48.2 | 34.3 | 32.0 | 23.1 | 34.6 | 30.3 | 32.0 | 32.4 |
| Method | Acc. | E2E Speedup |
|---|---|---|
| Qwen3-4B | 81.7 | |
| OmniKV | 80.0 | |
| Quest | 76.7 | |
| Quest ( Vortex ) | 75.0 | |
| SnapKV ( Low ) | 58.3 | |
| SnapKV ( Mid ) | 80.0 |
| Method | Resolved (%) |
|---|---|
| Vanilla | 23.7 |
| Quest | 21.3 |
| OmniKV | 22.3 |
| OmniKV ( Lag=4 ) | 25.0 |
| Method | Approx. | Time | E2E | E2E Out. |
|---|---|---|---|---|
| Max BS | (min) | Speedup | tok/s | |
| SWE-lite: GLM-4.7-Flash | ||||
| Vanilla | 40 | 98.1 | 302.0 | |
| Quest | 32 | 92.2 | 321.3 | |
| OmniKV | 40 | 62.4 | 474.5 | |
| SnapKV ( High ) | 52 | 82.0 | 361.2 | |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Transformer layer and query-head indices. | |
| Prompt length, current query position, and a key position. | |
| The causal history , including the current token. | |
| Query vector for head at position . | |
| Logical key and value used by query head at position . | |
| Query–key logit and normalized attention weight. |
| Method | Sink | Recent | Method budget | Other settings |
|---|---|---|---|---|
| SnapKV | 0 | 32 | 2,016 selected | 32 scoring window; pooling kernel 7 |
| Quest | – | – | 2,048 selected | 16-token pages; first two layers dense |
| OmniKV | 0 | 32 | 2,048 selected | Full layers: GLM ; Llama |
| H 2 O | – | – | 2,048 prefill, 2,048 decode | Prefill chunk eviction; recent ratio 0.5; requested prefill scoring window 128 |
| StreamingLLM (GLM) | 8 | 4,096 | – | – |
| PyramidKV (GLM) | 64 | 512 | 4,096 selected | Layer ratios to (linear); 32 scoring window |
| Method | Sink | Recent | Method budget | Other settings |
|---|---|---|---|---|
| SnapKV | 0/64 | 256 | 4,096 selected | 32 scoring window; pooling kernel 7 |
| Quest | 64 | 256 | 4,096 selected (4,416 total) | 16-token pages; first two layers dense |
| OmniKV | 64 | 256 | 4,096 selected | Full-attention layers: GLM ; Llama |
| H 2 O | – | – | 8,192 prefill; 4,096 configured decode | 128 prefill scoring window; recent ratio 0.5; decode eviction disabled |
| StreamingLLM (GLM) | 8 | 4,096 | – | – |
| PyramidKV (GLM) | 64 | 512 | 4,096 selected | Layer ratios to (linear); 32 scoring window |
| Method | Sink | Recent | Method budget | Other settings |
|---|---|---|---|---|
| OmniKV | 16 | 64 | 2,048 selected | Model-profile full-attention layers |
| Quest (SparseEngine) | 16 | 64 | 2,992 selected (3,072 total) | 16-token pages; first two layers dense |
| Quest (Vortex) | 16 | 64 | 2,048 selected (128 pages; 2,128 total) | First two layers dense |
| SnapKV | 16 | 64 | 4,096 selected (4,176 total) | 16 observation window; probability-based prefill scores; no full-attention layers; eviction every 1,024 decode steps |
| H 2 O | – | – | 4,096 prefill, 4,096 decode | 128 prefill scoring window; recent ratio 0.2; probability-based prefill scores; eviction every 1,024 decode steps |
| PyramidKV | 16 | 64 | 7,987 configured selected | 32 observation window; layer ratios to from layer 0; eviction every 1,024 decode steps |
| Method | Sink | Recent | Method budget | Prefix cache | Other settings |
| GLM-4.7-Flash | |||||
| SnapKV | 64 | 512 | 15,808 selected | Chain | 32 scoring window; probability-based prefill scores; no full-attention layers |
| H 2 O | – | – | 16,384 prefill; 8,192 decode | Chain | 128 prefill scoring window; recent ratio 0.5; FP32 logit scoring; decode eviction disabled |
| Quest | 64 | 512 | 1,472 selected (2,048 total) | Radix | 16-token pages; first two layers dense |
| OmniKV | 64 | 512 | 1,472 selected (2,048 total) | Radix | Model-profile full-attention layers |
| Qwen3-30B-A3B | |||||
| Parameter | SnapKV | Quest | OmniKV |
| sink_keep_tokens | 64 | 64 | 64 |
| recent_keep_tokens | 512 | 512 | 512 |
| decode_keep_tokens | 7,616 | 1,472 | 1,472 |
| full_attention_layers | – | – | auto |
| No. | Model | Attention Architecture | FFN Architecture |
|---|---|---|---|
| 1 | Qwen2.5 [ 40 ] | GQA | Dense |
| 2 | Qwen3 Dense [ 39 ] | GQA | Dense |
| 3 | Qwen3 MoE [ 39 ] | GQA | MoE |
| 4 | Qwen3.5 Dense | Hybrid linear attention | Dense |
| 5 | Qwen3.5 MoE | Hybrid linear attention | MoE |
| 6 | Qwen3.6 Dense | Hybrid linear attention | Dense |
| No. | Method | Sparsity Category |
|---|---|---|
| 1 | Quest [ 32 ] | Dynamic Sparsity |
| 2 | OmniKV [ 18 ] | Dynamic Sparsity |
| 3 | RetroInfer [ 10 ] | Dynamic Sparsity |
| 4 | FlashPrefill-v2 [ 14 ] | Dynamic Sparsity |
| 5 | StreamingLLM [ 37 ] | KV Eviction |
| 6 | SnapKV [ 24 ] | KV Eviction |
| Query-dependent retrieval | Physical KV eviction | Compressed KV representations | ||||||
| System | Quest | RetroInfer | SnapKV | Ada-SnapKV | H 2 O | Palu | LoRC | KIVI |
| SparseEngine (Ours) | ||||||||
| HiSparse | – | – | – | |||||
| Tangram | – | – | – | – | – | |||
| Vortex | – | – | – | |||||
| SPIN | – | – | – | |||||