Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.
Figures & tables
Figure 1: ( Left ) Clock Diffusion formulates semi-autoregressive continuous diffusion for discrete data, unlocking key drivers of practical LMs. ( Right ) ClockDLMs achieve SoTA diffusion language modeling, beating block discrete diffusion and setting a new standard for continuous LMs. Our Cache Grab sampler further pushes the frontier, exceeding AR accuracy on GSM8K.
Figure 2: Clock Diffusion training masks. Block : clean-to-clean is block causal, a noisy query attends to clean keys strictly preceding its own block, and noisy queries attend bidirectionally within a block. Sliding Window : we remove the outlined clean-to-clean entries, making that quadrant causal.
Figure 3: ( Left ) Cache Grab-Confidence. ( Right ) Cache Grab-SSD.
Table 4
Figure 4: Accuracy–efficiency trade-offs on GSM8K. Zero-shot test pass@ 1 versus average NFEs/seq for models trained on TinyGSM. (Left) Active window size S=32 . (Right) S=64 .
Figure 5: GSM8K Cache Grab. Cache Grab-Conf. achieves training-free speedups with minimal to no quality degradation. SSD unlocks improved performance & throughput.
NFEs/seq. ( ↓ )
ROUGE ( ↑ )
Mean ±SD
1
2
L
Autoregressive
73.1 ±13.4
35.9
14.8
25.0
Discrete Diffusion Baselines
MDLM
(full seq.)
180.0 ±0.0
39.5
17.1
26.1
( Sahoo et al., 2024a )
( S=64 )
128.0 ±0.0
40.1
18.1
27.2
( S=32 )
96.0 ±0.0
40.0
18.0
27.1
Table 3: Summarization metrics (ROUGE-1/2/L; ↑ ) on CNN/DailyMail. NFE values are means ± sample standard deviation. All models trained from scratch.
Figure 6: Cache Grab-Confidence on CNN/DailyMail. ROUGE-1 vs. NFEs/seq for ClockDLMs, sweeping the early-commit confidence threshold. Hollow markers use no early commit, i.e., rows in Table .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9
Figure 7: Fused caching and denoising attention masks for ( Left ) block and ( Right ) sliding window generation. Queries are the newly committed tokens xCs∣t followed by the active noisy tokens zsAs∣t ; keys are the cache KV followed by the same tokens. Committed tokens never attend to noisy ones, so the keys and values they write to the cache are exactly those a separate caching forward would produce. In this example, we assume the KV cache already has 4 tokens in it. For the block example, we are adding two newly decoded tokens and therefore starting to denoise 2 new samples from the prior on the subsequent denoising pass. For the sliding window example, we are adding one newly decoded token. The window on the subsequent pass consists of the token that was previously in the active window as well as a newly admitted sample from the prior.
Data
Tokenizer / vocabulary
gpt2 / 50,258
Context length L
1024
Packing
[EOS] per document; [BOS] / [EOS] per packed block
Validation split
Last 100,000 documents
Architecture
Backbone
DiT ( Peebles et al., 2023 ) with adaptive layer norm
Appendix
Table 4: Experimental configuration for OpenWebText.
Data
Tokenizer / vocabulary
HuggingFaceTB/SmolLM-135M / 49,153
Context length L
512 (right-padded, overlong examples filtered)
Packing
None; one example per sequence
Prompt / target
[BOS] ⊕ question ⊕ \n / code ⊕ [EOS]
Validation split
1% deterministic holdout of TinyGSM train
Architecture
Appendix
Table 5: Experimental configuration for TinyGSM training and GSM8K evaluation.
Data
Tokenizer / vocabulary
Qwen/Qwen3-0.6B-Base / 151,669
Context length L
768 ( 512 source +180 target, right-padded)
Target prefix
"Summary: "
Supervision
Target region only; padding neither noised nor scored
Architecture
Backbone
Qwen3-style decoder (architecture only; trained from scratch)
Appendix
Table 6: Experimental configuration for CNN/DailyMail.
Figure 8: MAUVE scores vs. NFEs/seq for sequences generated from models trained on OWT.
S=32
S=64
NFEs/seq. ( ↓ )
Acc. % ( ↑ )
NFEs/seq. ( ↓ )
Acc. % ( ↑ )
Mean ±SD
0 -shot, pass@ 1
Mean ±SD
0 -shot, pass@ 1
B-ClockDLM
Standard
190.1 ±47.3
54.3
220.1 ±44.3
49.1
+ Cache Grab-Conf.
( η=0.99 )
134.8 ±37.6
51.6
182.3 ±45.1
48.5
( η=0.95 )
117.5 ±39.8
53.8
161.8 ±47.0
47.2
Appendix
Table 7: Cache Grab-Confidence on GSM8K. Zero-shot pass@ 1 accuracy (Acc %; ↑ ) on the GSM8K test set and NFEs/seq for ClockDLMs and B-ClockDLMs trained on TinyGSM, sweeping the early-commit threshold η at N=S . NFE values are means ± sample standard deviations. Standard rows decode without early commits.
Diffusion loss only
Joint AR + diffusion loss
S
NFEs/seq. ( ↓ )
Acc. % ( ↑ )
NFEs/seq. ( ↓ )
Acc. % ( ↑ )
Δ
4
161.5 ±46.5
65.7
160.4 ±45.6
66.6
+ 0.9
8
165.0 ±45.8
63.8
164.5 ±45.6
67.2
+ 3.4
16
172.6 ±44.9
60.0
173.7 ±47.0
63.8
+ 3.8
32
188.1 ±44.8
59.6
187.4 ±44.4
63.5
+ 3.9
64
219.1 ±43.4
55.6
219.5 ±44.3
61.3
+ 5.7
Appendix
Table 8: Effect of joint AR and diffusion loss training. Zero-shot pass@ 1 on all 1,319 GSM8K test examples for Sliding Window ClockDLMs trained on TinyGSM, decoded with the standard diffusion sampler in both cases. Adding the AR loss leaves the denoising budget unchanged but improves accuracy at every active window size S . NFEs are means with ± sample standard deviations in subscripts.
Diffusion language models (DLMs) are an attractive alternative to autoregressive models because they promise sublinear-time, parallel generation, yet practical gains remain elusive as high-quality samples still demand hundreds of refinement steps. In continuous domains, consistency training along the probability-flow ODE is a popular recipe to accelerate diffusion. For discrete diffusion, no analogous sample-space ODE exists, making direct adaptation ill-defined. We argue that the right discrete substitute is the exact posterior bridge, the closed-form conditional law linking any two noise levels, which is available for broad corruptions including masked and uniform diffusion. Building on this observation, we introduce Multi-Path Discrete Consistency (MPDC), a new principle that trains a denoiser to be path-invariant in expectation across these stochastic bridges, and instantiate it as the Consistent Diffusion Language Model (CDLM), a single-stage training framework that does not require an already trained teacher model. Our CDLM objective recovers masked diffusion, continuous consistency models, and progressive or discrete distillation as analytic limits or empirical approximations of one common view. Empirically, CDLM establishes a new state of the art on both conditional and unconditional text-generation, consistently outperforming strong base discrete diffusion models and often even multi-stage distilled baselines across sampling budgets, with the largest gains in the few-step regime. Together, these results position CDLM as a principled and scalable foundation for the next generation of fast, high-fidelity discrete generative modeling.
Hasan Amin, Yuan Gao, Yaser Souri +4
Department of Computer Science, Purdue University · Microsoft
Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.
Jean-Marie Lemercier, Tomas Geffner, Karsten Kreis +3
While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than discrete approaches. To challenge this belief we revisit Plaid, a likelihood-based continuous diffusion language model (DLM), and construct RePlaid by aligning the architecture of Plaid with modern discrete DLMs. In this unified setting, we establish the first scaling law for continuous DLMs that rivals discrete DLMs: RePlaid exhibits a compute gap of only 20× compared to autoregressive models, outperforms Duo while using fewer parameters, and outperforms MDLM in the over-trained regime. We benchmark RePlaid against recent continuous DLMs: on OpenWebText, RePlaid achieves a new state-of-the-art PPL bound of 22.1 among continuous DLMs and superior generation quality. These results suggest that continuous diffusion, when trained via likelihood, is a highly competitive and scalable alternative to discrete DLMs. Moreover, we offer theoretical insights to understand the advantage of likelihood-based training. We show that optimizing the noise schedule to minimize the ELBO's variance naturally yields linear cross-entropy (information loss) over time. This evenly distributes denoising difficulty without any case-specific time reparameterization. In addition, we find that optimizing embeddings via likelihood creates structured geometries and drives the most significant likelihood gain.