Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a shared context's KV states depend on the tokens still being decoded, leaving directly reused caches stale and full recomputation necessary. We present ACache, a cross-request cache reuse mechanism for shared spans, or affixes, at any position: prefix, infix, or suffix. ACache measures the influence of affix tokens on the masked generation region to identify a small request-specific subset as Anchor Tokens, and recomputes only their KV states while reusing the remaining affix cache. Built on state-of-the-art intra-request caching mechanisms, ACache recovers most of the accuracy lost to direct affix-cache reuse on average when recomputing around 20% of affix tokens, and at that budget preserves more accuracy than selection criteria adapted from prior cross-request cache-reuse systems. We co-design ACache with a modern inference engine, whose attention reads each request's recomputed Anchor KV states alongside one affix cache shared across concurrent requests. Against the same system with only intra-request caching, ACache cuts recompute latency by up to 56.7%, translating to as much as 1.71× end-to-end throughput, while reducing peak KV cache memory by up to 45.8%.
Figures & tables
Figure 1: Cache reuse in AR LLMs and DLLMs. Left : AR LLMs reuse a shared prefix cache directly. Middle : in DLLMs, bidirectional attention makes that cache depend on the tokens still being decoded, leaving it stale. Right : ACache recomputes only the Anchor Tokens’ KV states.
Figure 2: Quality–throughput tradeoff. ACache’s Fast-dLLM-based system on LLaDA/GSM8K, 2-shot prefix, batch 16. Labels: Recompute ratios.
Figure 3: Anchor Token importance. Each panel shows a randomly selected GSM8K sample (Section 5 setup, 2-shot prompts). Color encodes mask-to-affix attention importance over affix tokens (vertical) and decoding steps (horizontal), highlighting a small subset at stable positions across steps.
Figure 4: System architecture of ACache. Left : requests reuse one precomputed affix cache as an infix (A), prefix (B), or suffix (C), selecting request-specific Anchor Tokens. Right : one request’s read-slot map, resolving each logical position to a private or shared KV slot.
Figure 5: Accuracy with ACache. Accuracy increases with the Anchor ratio (0: direct reuse; 1: full recomputation), except BABILong suffix, which declines overall.
Table 5: ACache’s criterion compared to adapted baselines, 1-shot. Accuracy (%) on Fast-dLLM. Bold marks the best value per column; underline, the second-best. BiCache † is the official implementation; Ratio is the mean equivalent Anchor ratio. Avg. omits BABILong suffix (Section 5.1 ).
Model
Dataset
1-shot
2-shot
4-shot
BiCache
ACache
BiCache
ACache
BiCache
ACache
LLaDA
GSM8K
96
99 ( 1.03× )
84
98 ( 1.18× )
71
92 ( 1.30× )
MBPP
192
213 ( 1.11× )
165
194 ( 1.18× )
118
155 ( 1.31× )
Dream
GSM8K
87
125 ( 1.44× )
93
120 ( 1.30× )
83
117 ( 1.41× )
MBPP
145
190 ( 1.31× )
136
183 ( 1.35× )
119
171 ( 1.44× )
Table 6: End-to-end throughput at batch size 1 under prefix reuse (tokens/s). Both run on Fast-dLLM: BiCache from its official implementation (Appendix C.3 ), ACache at Anchor ratio 0.2.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Affix
Shots
Model
Share
Jaccard similarity
Mass recall
Random
Fresh
ACache
Fresh
ACache
2nd
Last
2nd
Last
GSM8K
Prefix
1
LLaDA
62.7
11.3
71.6
56.8
38.9
98.5
95.8
76.7
Dream
71.8
11.3
78.8
57.4
54.4
97.8
89.5
89.3
2
LLaDA
62.6
11.2
73.4
52.5
35.2
98.5
95.1
91.4
Dream
73.4
11.2
84.9
58.2
62.8
99.4
93.6
94.3
Appendix
Table 7: Anchor stability at Anchor ratio 0.2. Fast-dLLM, full recomputation; all values in %, means over each test set. At each block boundary, a set is compared with the current top- K set: Random, Fresh (the first’s), or ACache’s Anchors (averaged). Share: the first top- K set’s importance.
Model
Dataset
Shot
Affix
Preparation
Forward
Top- K
Total
Forward share
LLaDA
GSM8K
1
213
0.49
50.05
0.16
50.69
98.7%
2
549
0.52
56.46
0.16
57.14
98.8%
4
981
0.57
66.02
0.18
66.76
98.9%
MBPP
1
178
0.55
54.70
0.16
55.40
98.7%
2
455
0.55
61.40
0.19
62.14
98.8%
4
1016
0.60
71.57
0.21
72.38
98.9%
Appendix
Table 8: Anchor selection cost per request (ms) under prefix reuse, batch size 1. Profiled on the runs of Table 2 (LLaDA) and Table 9 (Dream). Affix reports the shared span’s length in tokens.
Table 20: Keeping versus discarding the non-anchor affix cache, 1-shot. Accuracy (%) on Fast-dLLM. Anchors only selects the Anchor Tokens but discards the rest of the affix cache. Bold marks the better value per column. ACache rows repeat Table 5 ; as there, Avg. omits BABILong suffix.
Ratio
Selection
GSM8K
MBPP
BABILong
Avg.
Pfx.
Inf.
Suf.
Pfx.
Inf.
Suf.
Pfx.
Inf.
Suf.
0.2
Top
73.6
69.9
70.3
34.9
15.4
25.1
99.6
98.9
76.1
61.0
Random
73.2
69.0
69.0
34.1
11.1
18.6
99.6
80.4
84.5
56.9
Bottom
71.6
69.2
68.4
34.4
4.1
8.5
99.1
74.2
98.1
53.7
0.3
Top
74.1
71.0
70.8
34.8
24.7
26.8
99.7
98.9
78.6
62.6
Random
72.0
68.2
69.4
33.8
17.3
23.3
99.6
84.6
79.5
58.5
Appendix
Table 21: Top, random, and bottom Anchor selection, LLaDA, 1-shot. Accuracy (%) on Fast-dLLM; Top is ACache. Bold marks the best per column and ratio; Avg. omits BABILong suffix.
Model
Method
Ratio
Acc.
LLaDA
Direct reuse
0
19.8
ACache
0.2
59.1
0.3
65.4
BiCache †
0.50
66.9
Full recomputation
1
69.0
Dream
Direct reuse
0
78.5
Appendix
Table 22: System-prompt reuse on GSM8K: strict accuracy (%) on Fast-dLLM. As in BiCache, an answer counts only if the number after “####” matches. BiCache † is the official implementation; Ratio is its mean equivalent Anchor ratio. Bold marks the best reuse method per model; underline, the second-best.
Model
Batch 1
Batch 4
Batch 16
BiCache †
Fast-dLLM
ACache
Fast-dLLM
ACache
Fast-dLLM
ACache
LLaDA
62
76
91 ( 1.19× )
138
214 ( 1.55× )
164
290 ( 1.77× )
Dream
92
101
126 ( 1.26× )
167
273 ( 1.64× )
194
357 ( 1.84× )
Appendix
Table 23: System-prompt reuse on GSM8K: end-to-end throughput (tokens/s). Fast-dLLM is the full recomputation of Table 22 ; ACache uses Anchor ratio 0.2; BiCache † supports only batch size 1.
Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs). In DLMs, bidirectional attention means that updating any token dynamically alters the entire context and its corresponding KVs. Thus, existing caching techniques developed for LLMs, which assume that KVs remain invariant once computed, corrupt the shared prefix KVs. Our experiments show that applying these techniques to DLMs causes model accuracy to collapse to near zero. To unlock high-throughput DLM serving, we propose bidirectional prefix caching, BiCache, the first KV caching technique for shared prefixes in DLMs. BiCache is designed based on key observations from our comprehensive analysis: shared prefix KVs remain stable and reusable in shallow layers, while the depth of shallow layers depends on the fraction of shared prefix tokens in each request. Thus, BiCache dynamically identifies a safe layer depth for reusing shared prefix KVs and eliminates redundant computation. Evaluations demonstrate that BiCache significantly improves serving throughput by 36.3%-98.3% compared to existing techniques without accuracy collapse (only 0-1.8% difference).
Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the naive key--value (KV) cache behind fast autoregressive inference. Block diffusion restores caching by decoding block-by-block, and the block caches deployed on it so far are tied to attention: O(L)in memory and, if used as training-free retrofits, only an approximation of the model's computation. Both constraints can be overcome: sequence mixers that summarize finalized blocks into a reusable state support block caching, and the corresponding block-causal training objective makes the cache exact. We study this recipe at scale, pretraining three 3B block-diffusion denoisers (attention, mamba, and hybrid) on 300B tokens under one single-frontier objective and decoding all three through a single cached interface. Only the state-space cache is O(1) in sequence length: its memory and per-step latency stay constant at any context length, while an attention cache remains O(L). At 256k tokens (where attention has grown to 82GB and 29 ms/step), the Mamba cache delivers 4.3x lower latency, 11x less memory, and 2.6x higher single-stream throughput; and because that footprint is constant it scales with batch as well, reaching 14x the aggregate throughput, where attention cannot run beyond a single stream. The same linear-state bias lets the Mamba and hybrid backbones keep retrieving out to 8-16x their training length, whereas attention's retrieval collapses at 2x, at no measured quality cost.
Discrete diffusion language models (dLLMs) recover masked tokens in parallel, offering significant speedups over autoregressive (AR) generation. However, such promising frameworks face a fundamental architectural design dilemma: \ding{182} Adopting bidirectional attention achieves strong generation quality by allowing each position to access the full context, but is inherently incompatible with KV caching, limiting inference throughput in batch-serving scenarios; \ding{183} Conversely, causal attention enables efficient cached inference but loses all right-side context, substantially degrading generation quality. This paper introduces Bifocal dLLMs, a new paradigm that resolves this dilemma through \emph{asymmetric bidirectional context}. Analogous to bifocal lenses, we instantiate the paradigm as \textbf{R2LM} (Right-to-Left Mamba), which combines two complementary mechanisms: a) standard causal attention providing precise left-context with full KV cache compatibility, while b) a lightweight reverse Mamba SSM sidecar supplying compressed right-side context without breaking cacheability. Comprehensive experiments on continued pretraining of Qwen3-1.7B with 60B tokens demonstrate that R2LM achieves 2.4× to 12.9× higher throughput than bidirectional dLLMs and 1.9× to 2.9× speedup over AR baselines in batch serving through parallel decoding with KV caching, while exceeding the causal baseline on most benchmarks and surpassing the bidirectional dLLM on average.
Yuhang Chen, Xianfeng Wu, Jinhao Duan +11
University of North Carolina at Chapel Hill · Meta AI