cs.LGJul 14, 2026

AVQ-Attention: Adaptive Vector-Quantized Attention

Authors: Winfried van den doolPatrick ForréAmir HabibianYuki M. AsanoMax Welling

Organizations: QUVA Lab, University of Amsterdam, The Netherlands · AMLab, Informatics Institute, University of Amsterdam, The Netherlands · AI4Science Lab, University of Amsterdam, The Netherlands · Korteweg-de Vries Institute for Mathematics, University of Amsterdam, The Netherlands · Qualcomm AI Research, Amsterdam, The Netherlands · FunAI Lab, University of Technology Nuremberg, Germany

Abstract

The O(N2)\mathcal{O}(N^2) complexity of attention over NN tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces this to O(MN)\mathcal{O}(MN) by representing keys with MM codewords, but applies uniform codebook capacity regardless of where attention mass concentrates: high-attention regions of key space may be coarsely approximated while low-attention regions waste representational capacity. We propose Adaptive Vector-Quantized (AVQ) Attention, which adaptively allocates codebook capacity based on attention importance. Starting from a small set of codewords, our method identifies the most important codes during the forward pass and refines them with pre-learned child codewords, achieving fine-grained quantization where it matters most while maintaining coarse quantization elsewhere. We develop an implementation using custom Triton kernels that enables the full adaptive refinement process, including importance scoring, child codeword insertion, and parent contribution replacement, to be carried out within the tiled computation paradigm of Flash Attention with minimal overhead. Our approach maintains O(MN)\mathcal{O}(MN) complexity while achieving improved accuracy-efficiency trade-offs compared to fixed-codebook VQ-attention.

Explore similar work

CardsList
  1. KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

    Sep 21, 2026Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon