cs.LGSep 28, 2026

SANTA++: Sampling Attention through Representative Keys

Authors: Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari

Organizations: Department of Electrical and Computer Engineering University of California, Santa Barbara, CA, USA · Flucta, San Francisco, CA, USA

Abstract

Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusion probability. This importance sampling correction estimates attention over the full cache, with a sampling budget that lets us trade memory reads for accuracy. Remarkably, with 32 or 64 sampled teams, SANTA++ uses 16% to 22% of dense attention's KV reads and retains 94% to 99% of the dense-attention baseline's scores on LongBench v2 and HELMET's retrieval-augmented generation subset, and 85% to 91% on RULER, with Qwen2.5-7B-Instruct at 32K context. With 31 sampled teams, our GPU implementation delivers a 1.69×1.69\times attention speedup over the dense FlashAttention baseline at 32K context. By reducing the number of cache entries read, SANTA++ in principle complements architectures with compressed KV representations, such as multi-head latent attention. Our kernels are available at: https://github.com/OPUSLab/santapp-kernel-demo.git.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Stochastic Sparse Attention for Memory-Bound Inference

    May 3, 2026Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5Dynamic Sparse AttentionAutoregressive Decoding

  2. Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

    Jun 20, 2026Xin GaoTransformer ArchitecturesReasoning Benchmark

  3. Simplified Sparse Attention via Gist Tokens

    Apr 22, 2026Yuzhen Mao, Michael Y. Li, Emily B. FoxDynamic Sparse AttentionEfficient Long-Context Inference