PageWeaver: KV-Guided Query Unions for Sparse Attention
Organizations: SenseTime
Abstract
Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs. A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost. With FP8 KV throughout, the H200 Union8 implementation achieves a 1.70x geometric-mean complete-call speedup over the measured FlashInfer path on six captures. Online regrouping further lowers latency by 3.26-7.66% on five selected 64K-context captures. Whole-model prefill throughput is 7.88-14.36% above the tested native path; the incremental regrouping benefit is smaller, with observed median gains of 0.47-0.73% at 32K/64K and regressions at 8K. A B300 comparison identifies cases where preparation cost and a stronger native kernel remove the advantage. These results separate execution-group reuse from the complete cost of exploiting it online.
Figures & tables
| Study | Workloads and repetitions |
|---|---|
| Base execution | Six captures, one request; L3/L32/L59, 0/16K prefixes. Seven rounds 20 replays. |
| Online grouping | Five captures of one constructed 64K input: L39/r0–3, L43/r1; plus one 16K L51 control. Eleven rounds 50. |
| H200 service | Five frozen prompts per 8K/32K/64K bucket; concurrency 1/4; three measured blocks per arm/cell. |
| B300 auxiliary | Eight TP4/rank0 captures: L3/L20/L32/L59 at two chunks. Seven rotated graph/eager rounds. |
| Backend | NRMS (%) |
|---|---|
| Union8 | 2.246–3.967 |
| FlashInfer | 2.288–4.077 |
| Native SM90 | 6.258–9.120 |