cs.LGOct 4, 2026

Mask-Guided KV Cache Eviction in Block Diffusion Language Models

Authors: Gleb Molodtsov, Ekaterina Alimaskina, Evgeny Uskov, Artur Zagitov, Aleksandr Beznosikov

Abstract

Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for denoising the current block (selection) and which to keep in memory for future blocks (eviction). We propose MaskAhead, a training-free method that solves both tasks with a single mask-query-based ranking mechanism. Current-block masks guide selection, while probes of upcoming masked blocks guide eviction. Both rank KV entries by their estimated contribution to the attention output. Our quantized variant, Q-MaskAhead, computes selection and attention directly from low-bit KV, largely preserving the selected entries. Experiments on Fast-dLLM-v2, DreamReasoner, and LLaDA2.0-mini cover long-generation reasoning, long-prompt question answering, and needle-in-a-haystack retrieval. On long-prompt QA, MaskAhead reduces KV memory by 9.5×9.5\times on average with a 1.2-point mean F1 loss relative to dense inference. Q-MaskAhead increases the reduction to 20.1×20.1\times with a 2.3-point mean F1 loss. In a batch-32 systems profile, MaskAhead achieves 1.23×1.23\times end-to-end and 1.68×1.68\times decode-stage speedups over dense inference.

Explore similar work

CardsList
  1. Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

    Sep 9, 2026Vaibhav Singh, Pierre-André Noël, Torsten Scholak +2Blockwise Diffusion Language ModelsDiffusion Language Model Inference

  2. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

    May 10, 2026Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1Self-AttentionKV Caching

  3. HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval

    Jun 19, 2026Omin Kwon, Doyeon Kim, Jongseok Park +3KV-Cache OffloadingDiffusion Model Serving