cs.CLJun 26, 2026

Affix Cache for Diffusion Large Language Models

Authors: Kaihua Liang, An Zhong, Xin Tan, Zafar Ayyub Qazi, Hong Xu, Jian Weng, Marco Canini

Organizations: KAUST · CUHK · LUMS

Abstract

Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a shared context's KV states depend on the tokens still being decoded, leaving directly reused caches stale and full recomputation necessary. We present ACache, a cross-request cache reuse mechanism for shared spans, or affixes, at any position: prefix, infix, or suffix. ACache measures the influence of affix tokens on the masked generation region to identify a small request-specific subset as Anchor Tokens, and recomputes only their KV states while reusing the remaining affix cache. Built on state-of-the-art intra-request caching mechanisms, ACache recovers most of the accuracy lost to direct affix-cache reuse on average when recomputing around 20% of affix tokens, and at that budget preserves more accuracy than selection criteria adapted from prior cross-request cache-reuse systems. We co-design ACache with a modern inference engine, whose attention reads each request's recomputed Anchor KV states alongside one affix cache shared across concurrent requests. Against the same system with only intra-request caching, ACache cuts recompute latency by up to 56.7%, translating to as much as 1.71×\times end-to-end throughput, while reducing peak KV cache memory by up to 45.8%.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Enabling KV Caching of Shared Prefix for Diffusion Language Models

    May 26, 2026Younghun Go, Jaehoon Han, Changyong Shin +2CacheDiffusion Language Models

  2. Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

    Sep 14, 2026Vaibhav Singh, Pierre-André Noël, Torsten Scholak +2Diffusion Language ModelsCache

  3. Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation

    Jun 26, 2026Yuhang Chen, Xianfeng Wu, Jinhao Duan +11Diffusion Language ModelsBidirectional