cs.LGSep 30, 2026

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

Authors: Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu

Organizations: Amazon AGI

Abstract

Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the QKTQK^T scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to 1.25×1.25\times the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to 1.71×1.71\times with comparable accuracy.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference

    Feb 8, 2026Jitai Hao, Qiang Huang, Yaowei Wang +2Efficient Long-Context InferenceKey-Value Cache Compression

  2. SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

    Jun 30, 2026Amirhossein Abaskohi, Giuseppe Carenini, Peter West +1Key-Value Cache CompressionKey-Value Cache