Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This is implemented through a two-stage architecture consisting of a relatively expensive anchor network that generates the latent cache state and a lightweight denoising network that intelligently combines the cached latent state with the current state at each reverse step using a fusion module. This gives anchoring a latent-space caching interpretation: the anchor network is evaluated periodically, while its cached representation is reused across multiple reverse steps. We instantiate this framework as TADM:Post-train, which time-anchorizes pretrained DLMs, and TADM:Pretraining, which learns time-based anchors during pretraining. Applied to DiffusionGemma-26B, TADM:Post-train improves throughput by approximately 49% to 79% on several math, code, and STEM benchmarks (GSM8K, AIME26, GPQA-Diamond, LiveCodeBench-v6, HumanEval, MMLU-Pro). TADM:Pretraining reduces Transformer-layer computation by up to 38% relative to a standard single-stage DLM, achieves up to 73% higher measured throughput than ADLM.
Figures & tables
Figure 1: Time-Anchored Diffusion Model (TADM). The expensive anchor pathway is evaluated once every K reverse steps to produce a latent anchor ht′ , which is cached and reused. At each subsequent step, the lightweight shared network encodes the evolving state zt , and a gated fusion module combines its current representation ct with the cached anchor before denoising. TADM therefore replaces repeated evaluation of the expensive anchor network with learned latent-space reuse.
Benchmark
Method
K
Acc. (%)
Δ Acc. (pp)
Tok./s
Throughput
Compute ( × Base)
GSM8K
DiffusionGemma
-
94.79
–
33.80
1.00 ×
1.00 ×
GSM8K
TADM
1
94.79
0.00
35.72
1.06 ×
1.00 ×
GSM8K
TADM
2
94.79
0.00
51.51
1.52 ×
0.67 ×
GSM8K
TADM
3
94.95
+0.16
59.53
1.76 ×
0.56 ×
GPQA-D
DiffusionGemma
-
67.00
–
60.31
1.00 ×
1.00 ×
GPQA-D
TADM
1
67.00
0.00
55.30
0.92 ×
1.00 ×
Table 1: TADM:Post-train performance. Results are averaged over three seeds across math, code, and STEM benchmarks and compared with the DiffusionGemma baseline under the same fixed-step evaluation setting. Increasing the anchor refresh interval enables progressively greater reuse of the cached latent representation, yielding substantial inference speedups. TADM preserves essentially the same task accuracy across benchmarks while achieving up to 1.79× higher generation throughput.
Method
MAUVE ( ↑ )
Gen PPL ( ↓ )
Entropy ( ↑ )
Compute ( ↓ )
Throughput ( ↑ )
Data
1.00
14.8
5.44
-
-
AR (T=1024)
0.760
12.1
5.22
-
-
T=2048
T=4096
T=2048
T=4096
T=2048
T=4096
T=2048
T=4096
T=2048
T=4096
SEDD (absorb, 170M)
0.008
0.009
103.2
102.5
5.61
5.61
24576 (1.00x)
49152 (1.00x)
37.47 (1.54x)
18.79 (1.34x)
MDLM (170M)
0.037
0.035
51.3
50.9
5.46
5.45
24576 (1.00x)
49152 (1.00x)
63.00 (2.59x)
46.39 (3.31x)
MDLM+FB (170M)
0.197
0.243
28.6
22.8
5.28
5.18
24576 (1.00x)
49152 (1.00x)
37.01 (1.52x)
25.28 (1.80x)
Table 2: We compare TADM diffusion baselines across sampling budgets. For TADM, we use refresh intervals K={2,2,2,8,4,4} for T={128,256,512,1024,2048,4096} , respectively. We report MAUVE, generative perplexity (Gen PPL), entropy, Transformer-layer evaluations normalized to MDLM over 5000 samples, and measured generation throughput over 20 generations of length 1,024. TADM periodically reuses its cached anchor representation, reducing Transformer computation while maintaining competitive generation quality.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Lambada
PTB
Wikitext
LM1B
AG News
PubMed
ArXiv
AR
51.28
82.05
25.75
51.25
52.09
49.01
41.73
AR+Diffusion
BD3-LM ( L′=4 )
50.03
96.81
31.31
60.88
61.67
42.52
39.20
Diffusion
SEDD
49.86
100.09
34.28
68.20
62.09
44.53
38.38
MDLM
47.52
95.26
32.83
67.01
61.15
41.89
37.37
Appendix
Table 3: Zero-shot validation perplexities ( ↓ ) on OWT-trained models with 1,024 NFEs. TADM:Pretraining † is the γ=0 variant and TADM:Pretraining ‡ is γ=3e−3 .
Method
K
T=128
T=256
T=512
Tok/s ↑
Gen PPL ↓
Ent.
Tok/s ↑
Gen PPL ↓
Ent.
Tok/s ↑
Gen PPL ↓
Ent.
ADLM
1
320.62
57.29
5.51
161.31
41.88
5.46
81.44
31.63
5.33
2
404.51
78.98
5.59
202.45
60.78
5.52
102.17
48.01
5.46
4
462.16
194.00
5.66
231.85
106.46
5.58
117.30
76.38
5.44
8
498.35
680.71
5.86
250.77
245.08
5.65
126.79
150.22
5.47
TADM †
1
347.69
42.25
5.448
175.25
27.97
5.347
88.11
20.46
5.252
Appendix
Table 5: Effect of stale-anchor reuse in ADLM and TADM:Pretraining. We compare anchor refresh intervals K∈{1,2,4,8} , where K=1 recomputes the anchor at every reverse step. We report throughput (Tok/s), generative perplexity (Gen PPL), and token entropy.
Accuracy (%)
Corruption (%)
Tok./s (speedup)
Benchmark
K
Baseline
Ours
Baseline
Ours
Baseline
Ours
AIME26
1
47.78
47.78
11.11
11.11
62.79 (1.00 × )
57.48 (0.92 × )
2
44.44
46.67
15.56
14.44
90.67 (1.44 × )
83.41 (1.33 × )
3
44.44
48.89
37.78
20.00
105.67 (1.68 × )
93.50 (1.49 × )
GSM8K
1
94.79
94.79
0.08
0.08
33.80 (1.00 × )
35.72 (1.06 × )
2
94.69
94.79
0.13
0.05
51.82 (1.53 × )
51.51 (1.52 × )
Appendix
Table 6: Naive stale-anchor reuse versus TADM:Post-train. Results are averaged over seeds {0,1,2} with adaptive stopping disabled. The original cache-age notation k={0,1,2} is reported here as anchor refresh intervals K={1,2,3} . “Baseline” denotes naive reuse with the original DiffusionGemma weights and no fusion modules, while “Ours” denotes TADM:Post-train. Speed is reported as tokens/s, with speedup in parentheses relative to the corresponding baseline K=1 throughput. Corruption is the percentage of generations flagged as incoherent or containing a repetition loop.
Diffusion language models (DLMs) generate text through iterative denoising, but inference requires full-sequence attention at every iteration, resulting in substantial redundant computation on masked tokens. Block-wise diffusion can reduce this cost, yet it typically relies on retraining and constrained update orders, limiting its direct applicability to pretrained DLMs. Our token-level analysis reveals pronounced structural locality in DLM inference. Decoding is driven by a small set of prefix-localized active tokens; the influence of distant undecoded context diminishes rapidly, and decoded tokens exhibit stage-wise temporal stability, enabling reuse of intermediate representations except for a brief post-decode transient. Motivated by these observations, we propose \textbf{\placeholder}\footnote{The source code is available at https://github.com/vhicrgit/Window-Diffusion.}, a window-based token pruning and caching method for inference. We maintain a local computation window that slides rightward as denoising progresses, and partition undecoded tokens into: (i) \textit{active tokens} that are computed online, (ii) \textit{buffer tokens} whose KV states are cached and periodically refreshed, and (iii) \textit{far-field tokens} that are pruned outside the window. Computation is restricted to active and buffer tokens within the window, while far-field tokens are omitted at each stage. Experiments on LLaDA and Dream show that, under matched compute budgets, our method achieves up to 99× inference speedup while largely preserving generation performance.
Fengrui Zuo, Zhiwei Ke, Yiming Liu +3
School of Computer Science and Technology, University of Science and Technology of China · Suzhou Institute of Advanced Research, University of Science and Technology of China
Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.
Jean-Marie Lemercier, Tomas Geffner, Karsten Kreis +3
Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes inference costly. Every denoising update alters the global context, forcing both prompt and response states to be recomputed even though only response tokens are revisable. Key-value (KV) caching could reduce this cost, yet conventional caching assumes immutable historical states and is therefore difficult to reconcile with rollback. In this paper, we introduce Adaptive Reuse of Cached Hidden States for Efficient Rollback (Archer), a training-free KV caching method for rollback-capable DLMs. Archer asymmetrically keeps the mutable response synchronized with the current hypothesis while reusing prompt K/V within a bounded state neighborhood. Although prompt representations also change under bidirectional attention, their token identities remain fixed; bounded reuse therefore amortizes repeated prompt computation without caching mutable response states. It also delays feedback from tentative tokens, reducing premature reinforcement of transient high-confidence errors and giving rollback more opportunity to correct them. Our analysis characterizes prompt reuse as a reversibility-aligned cache boundary, bounds its state-dependent approximation error, and gives a decoder-margin condition for preserving full-refresh decisions. Existing DLM acceleration often trades quality for speed. Archer shifts this frontier, attaining the best mean performance of 33.63% together with a 2.57x mean speedup on the main suite. Across evaluated settings, it improves Pass@1 by up to 3.05 points and reaches up to 2.95x speedup. Controlled analyses connect the quality gain to delayed prompt feedback and validate state-aware refresh. Our code is available at https://github.com/Hxnng/Archer.
Xuning He, Zinan Sheng, Yongding Tao +4
School of Computer Science, Shanghai Jiao Tong University · College of Artificial Intelligence, Nankai University · School of Computer Science, Peking University