Discrete diffusion models and flow matching have emerged as powerful frameworks for generative modeling over discrete state spaces, yet efficient few-step generation remains a fundamental challenge. In this work, we introduce the Discrete Average Generator, a principled extension of MeanFlow to Continuous-Time Markov Chains (CTMCs). Analogously to how MeanFlow defines an average velocity field over a time interval in continuous spaces, we define an average generator as the normalized increment of the transition kernel over a time interval. We show that this average generator satisfies a self-consistency identity, which provides the foundation for our training objective. We further develop training strategies that align with the standard training paradigm of diffusion language models while keeping the resulting objective tractable. When projected onto per-coordinate marginals, the self-consistency identity admits a closed-form expression, enabling efficient training and inference. In Potts model simulations, our objective reduces the total variation distance of the K-step sampler by up to 67%. On OpenWebText, our method achieves the lowest generative perplexity among the evaluated methods for 8 to 64 sampling steps while enabling a 16× acceleration, and achieves comparable performance to existing methods on ImageNet.
Figures & tables
Property
MeanFlow
Discrete Average Generator
State Space
Continuous domain ( Rd )
Discrete space ( SD ) via CTMCs
Instantaneous Measure
Velocity field v(xt,t)
Transition rate generator Qt(x,z)
Average Measure
Average Velocity: u(xt,t,r)=r−t1∫trv(xs,s)ds
Average Generator: Ut,r=r−tPt→r−I
Material Derivative
Spatial gradient used: dtdu=∂tu+v⋅∇xu
Dynkin’s formula used: DtDU=∂tU+QtU
Fixed-Point Identity
u=v−(t−r)(∂tu+v⋅∇xu)
Ut,r=Qt−(t−r)(∂tUt,r+QtUt,r)
Stable Target Parameterization
x -prediction (predicting clean data x1 )
p~ -prediction (predicting a per-coordinate distribution p~t,r(⋅∣xt) )
Table 1: Conceptual and mathematical comparison between MeanFlow and the proposed Discrete Average Generator.
Figure 2
Figure 3: Generative perplexity and entropy on OpenWebText. The dashed horizontal line indicates the PPL of Di4C + ReDi 1 at K=1024 (18.44). Values in Table 3 .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
K
2
3
4
6
8
12
16
24
32
Baseline
0.6020
0.3707
0.2456
0.1351
0.0907
0.0508
0.0356
0.0225
0.0136
Ours
0.5381
0.3170
0.1934
0.0957
0.0517
0.0181
0.0116
0.0080
0.0053
Appendix
Table 2: Total variation Potts model simulations. Mean over three seeds. These are the values plotted in Figure 2 .
Sampling steps K
Model
4
8
16
32
64
Generative perplexity
DUO+DCD
97.22
72.18
54.82
46.05
42.38
ReDi 1
79.35
48.53
37.13
31.21
29.12
ReDi 2
69.01
45.11
36.42
31.56
29.85
ReDi 3
53.24
36.33
30.34
27.75
26.78
Appendix
Table 3: OpenWebText results for Figure 3 : generative perplexity measured with LLaMA 3.1-8B and entropy, as a function of the number of sampling steps K . Baseline numbers are taken from Yoo et al. (2025) . Best perplexity per column in bold.
Model
FID ↓
IS ↑
MaskGIT ( Chang et al., 2022 )
10.90
184
SDTT ( Deschenaux and Gulcehre, 2025 )
8.97
205
Di4C ( Hayakawa et al., 2025 )
6.20
216
ReDi 1 ( Yoo et al., 2025 )
7.58
228
ReDi 2 ( Yoo et al., 2025 )
7.86
240
Ours
7.81
250
Appendix
Table 4: ImageNet 256×256 with four network evaluations. Every method is reported at the best sampling configuration found by a sweep. Best value per column in bold.
Figure 4: Ablation on the classifier-free guidance weight w used to sample the training pairs, evaluated at 4 - and 8 -step generation on ImageNet 256×256 (50k samples). (a) Precision and recall vs. w . (b) FID vs. w , with the best configuration at each step budget annotated. (c) Inception score and density vs. w .
Figure 5: Qualitative effect of the CFG weight w used to sample the training pairs, at 4 -step generation on ImageNet 256×256 , for nine classes: goldfish, ostrich, hourglass, racer, volcano, goldfinch, sea anemone, cheeseburger and spotted salamander. The random seed is held fixed per class across w∈{1.0,1.5,2.0,3.0,4.0,5.0} , so only the effect of w varies; the optimal 4 -step setting ( w=3.0 , FID 7.81 ) is highlighted.
Figure 6: 4 -step generation results on ImageNet (classes goldfinch, sea anemone, cheeseburger, spotted salamander), two samples per class and model. Our method (bottom row, FID 7.81 ) is appended below the baselines.
Diffusion language models (DLMs) are an attractive alternative to autoregressive models because they promise sublinear-time, parallel generation, yet practical gains remain elusive as high-quality samples still demand hundreds of refinement steps. In continuous domains, consistency training along the probability-flow ODE is a popular recipe to accelerate diffusion. For discrete diffusion, no analogous sample-space ODE exists, making direct adaptation ill-defined. We argue that the right discrete substitute is the exact posterior bridge, the closed-form conditional law linking any two noise levels, which is available for broad corruptions including masked and uniform diffusion. Building on this observation, we introduce Multi-Path Discrete Consistency (MPDC), a new principle that trains a denoiser to be path-invariant in expectation across these stochastic bridges, and instantiate it as the Consistent Diffusion Language Model (CDLM), a single-stage training framework that does not require an already trained teacher model. Our CDLM objective recovers masked diffusion, continuous consistency models, and progressive or discrete distillation as analytic limits or empirical approximations of one common view. Empirically, CDLM establishes a new state of the art on both conditional and unconditional text-generation, consistently outperforming strong base discrete diffusion models and often even multi-stage distilled baselines across sampling budgets, with the largest gains in the few-step regime. Together, these results position CDLM as a principled and scalable foundation for the next generation of fast, high-fidelity discrete generative modeling.
Hasan Amin, Yuan Gao, Yaser Souri +4
Department of Computer Science, Purdue University · Microsoft
Discrete diffusion promises orders-of-magnitude faster generation than autoregressive (AR) models for sequential discrete data, yet its full potential of few-step generation has remained out of reach due to a fundamental structural limitation. The conditional-independence assumption underlying current discrete diffusion models introduces a systematic parallelization bias that compounds with the number of tokens unmasked per step, becoming severe in the few-step regime that fast generation requires. We address this with the first framework for explicit joint distribution modeling in discrete diffusion via tensor decomposition, which represents the conditional clean distribution as a low-rank tensor with controllable expressivity. The framework supports both Canonical Polyadic (CPD) and Tensor-Train (TTD) decompositions, and we identify a structural bias of TTD toward dependencies between nearby tokens, formalized through Oseledets' theorem relating TT-rank to unfolding-matrix rank, which is well-suited to sequential data such as natural language and line notations for molecular data. To enable efficient generation, we present an iterative marginal inference procedure with specialization for predetermined position schedules. Our framework integrates into pretrained MDMs through lightweight fine-tuning, yielding substantial improvements in few-step generation at a fraction of the cost of training from scratch.
Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.