Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emph{Branching Latent Diffusion (BLD)}, which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a 16× reduction. BLD combines latent compression with \emph{branching token realization}, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than 80× and increases throughput by more than 6× relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than 6× higher throughput and more than 4× lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.
Figures & tables
Figure 1: Overview of BLD. A fusion encoder compresses L tokens into K≪L block latents. The latent sequence is partitioned into groups of G blocks and generated autoregressively across groups, with each group conditioned on the previously generated latents. A branching decoder then expands the complete latent sequence into K token spans and decodes them in parallel.
Model
Params
States
Gen-PPL ↓
Entropy ↑
MAUVE ↑
TFLOPs ↓
OpenWebText
–
–
15.3
5.32
0.95
–
AR
730M
1024
26.5
5.37
0.92
1.6
MDLM
170M
1024
35.2
5.30
0.79
298.4
BD3-LM, L′=4
170M
1024
23.4
5.27
0.82
17.9
BD3-LM, L′=8
170M
1024
28.7
5.31
0.86
17.3
BD3-LM, L′=16
170M
1024
32.4
5.33
0.85
18.7
Table 1: Generation quality and computation for approximately 1024-token sequences. Params counts the parameters used at generation time (prior and branching decoder for BLD; the fusion encoder is not used during generation). States is the number of positions processed by the denoiser. ELF values use the released checkpoints and their default sampling settings; TFLOPs are measured by us. † Fully parallel variant with a separately tuned recipe (Appendix C.1 ).
Model
Precision
Latency (s) ↓
Throughput (doc/s) ↑
MDLM
bf16
30.15
0.4
BD3-LM, L′=4
fp32
16.77
0.9
BD3-LM, L′=8
fp32
15.37
1.0
BD3-LM, L′=16
fp32
14.89
1.0
ELF-B, 32 steps
bf16
0.46
18.2
ELF-B, 64 steps
bf16
0.90
9.5
Table 2: End-to-end runtime on a single NVIDIA H100 80GB GPU. Latency is the median over five post-warm-up runs at batch size 1, and throughput is measured at batch size 32. Both include sampling and token decoding. All models use eager execution at the listed precision; details are provided in Appendix C.2 . † Fully parallel variant (Appendix C.1 ).
G
Gen-PPL ↓
Entropy ↑
Latency (s) ↓
TFLOPs ↓
1
31.7
5.22
4.93
2.23
4
24.7
5.21
3.52
1.31
8
26.6
5.17
1.84
1.10
16
27.7
5.03
1.21
1.00
Table 3: Effect of the group size G . All models use the same compressed latent representation, branching decoder, grouped prior architecture, and training recipe; only G varies. Latency is measured at batch size 1 on an H200.
Decoder
Gen-PPL ↓
Entropy ↑
MAUVE ↑
Rec. ↓
Not adapted
127
5.34
0.67
0.042
Adapted on G=K
26.6
5.17
0.82
0.030
Adapted on G=K , then G=8
27.1
5.15
0.78
0.027
Table 4: Effect of decoder adaptation with the same G=8 prior. Rec. denotes reconstruction NLL from exact encoder latents.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Tokens per latent
Clean reconstruction
Midpoint recovery
1
1.000
0.925
2
0.999
0.791
4
0.952
0.455
Appendix
Table 5: Token accuracy under increasing latent compression. Clean reconstruction decodes the uncorrupted latent. Midpoint recovery measures recovery of the source text from a latent corrupted to the midpoint of the diffusion trajectory. Both columns report token accuracy on 64 validation sequences using the final five-epoch checkpoint; midpoint recovery decodes a one-step clean-state prediction at t0=0.5 with self-conditioning guidance scale 1.
Rank
Retained variance
Oracle PPL
Generated PPL
Transport gap
512
1.000
30.56
554.19
2.898
256
0.879
31.84
441.01
2.628
224
0.834
36.34
476.70
2.574
192
0.780
55.66
393.98
1.957
128
0.635
199.38
259.11
0.262
Appendix
Table 6: PCA rate scan on a frozen K=64 , L=1024 block-latent autoencoder. Each prior is trained for 5,000 steps; generation uses 256 paired samples, greedy decoding, and a 256-token GPT-2 scoring window. Oracle PPL scores decoded real latents after projection; generated PPL scores prior samples. The transport gap is log(generatedPPL/oraclePPL) . These development-stage scores are not directly comparable with Table 1 .
a
Exact match with A
PPL
0.125
99.9%
30.6
0.250
97.8%
33.6
0.375
67.5%
141.6
0.500
19.0%
680.8
Appendix
Table 7: Decoding interpolated real-document latents z(a)=(1−a)zA+azB with a frozen encoder and decoder. Exact match measures agreement with document A ; PPL is GPT-2-scored on a 256-token window. The probe uses 64 held-out documents and is separate from the main generation evaluation.
Prior
Gen-PPL
Entropy
Distinct-1
Block AR, noise scale 0.6
19.3
5.35
0.252
Fully parallel, noise scale 0.7
28.1
5.30
0.220
Appendix
Table 8: Development-stage block-autoregressive (AR) latent prior and fully parallel diffusion control on the same frozen autoencoder and decoder. Gen-PPL is scored by GPT-2 Small on 256-token windows from 256 samples, rather than the full-document protocol of Table 1 . The AR prior uses a 400M backbone and 53M flow head with 50 SDE steps per block; the parallel prior has 684M parameters and uses 50 steps over the whole sequence.
Fine-tuning
Noise scale
Gen-PPL
Entropy
Distinct-1
Loop rate
Control
0.70
25.83
5.304
0.220
0.86
R1 ( p=0.5 )
0.60
26.77
5.43
0.228
0.82
R3 ( p=1 )
0.60
24.65
5.395
0.220
0.90
R4 ( p=1 , tmin=0.7 )
0.70
23.10
5.195
0.210
–
Appendix
Table 9: Development-stage rollout fine-tuning of the fully parallel prior, with 50-step sampling. Gen-PPL is scored by GPT-2 Large on 256-token windows; loop rate is measured on 50 inspected full texts per point. Noise scales and resulting entropies differ across rows; these scores are not the full-document scores in Table 1 .
Model
fp32
bf16 autocast
bf16 weights
BLD, G=8
2.67 / 9.4
3.76 / 5.8
3.46 / 5.9
BLD, G=K
1.16 / 13.5
1.40 / 13.2
1.33 / 13.4
AR
12.57 / 1.5
16.15 / 0.56
14.29 / 0.58
Appendix
Table 10: Runtime under different precisions, using the same GPU and measurement protocol as Table 2 . Entries report latency (s) / throughput (doc/s). For bf16 weights, parameters are cast to bf16 once and inference also uses bf16 autocast.
share of blocks
longest span
boundary position (%)
Documents
docs
mi (mean ± std)
mi=16
mi≤4
median / p95 / max
word start
after sentence end
Encoder (chunker), real documents
Held-out, full length
330
16.00±0.18
96.7%
0.0%
17 / 17 / 17
65.8 (65.3)
3.9 (3.7)
Training, full length
1000
16.00±0.31
96.5%
< 0.1%
17 / 17 / 31
64.8 (64.8)
3.6 (3.7)
Held-out, all lengths
1000
14.98±3.17
84.9%
6.0%
16 / 22 / 24
69.0 (69.2)
3.8 (3.6)
Decoder (length head), generated documents
Appendix
Table 11: Encoder block lengths and predicted decoder branch lengths ( L=1024 , K=64 ). Full length: documents with at least L tokens. All lengths: held-out documents with at least K tokens, truncated to L and right-padded as in training (median length 983; 2.5% reach L ). Longest span denotes maximi per document. Parentheses give boundary-position base rates over all token positions.
Chunker
PR
k99
Recon. CE
Gen-PPL
Entropy
Distinct-1
Loop rate
Dynamic
9.7
180
0.156
105.4
5.831
≈ 0.305
0.12
Uniform
33.3
243
0.041
343.2
6.049
≈ 0.305
0.38
Appendix
Table 12: Paired dynamic-versus-uniform chunking diagnostic on a reduced development stack. PR and k99 characterize the autoencoder latent spectrum; reconstruction CE uses real latents. Generation points have approximately matched distinctness, but their tested entropy ranges do not overlap, so the PPL entries are not an iso-entropy comparison and must not be compared with Table 1 .
Model
Length
States
Gen-PPL ↓
Entropy ↑
MAUVE ↑
TFLOPs ↓
Main-text evaluation: approximately 1024 tokens
OpenWebText
–
–
15.3
5.32
0.95
–
AR
1024
1024
26.5
5.37
0.92
1.6
ELF-B, 32 steps
1024
1024
23.8
5.15
0.80
7.2
BLD, G=8
1024
64
26.6
5.17
0.82
1.1
BLD, G=K †
1024
64
24.6
5.08
0.77
4.8
Appendix
Table 13: Cosmos measurements and main-text results under different evaluation lengths. Baseline and BLD scores use approximately 1024-token sequences; Cosmos scores use prefixes of at most 128 or 256 GPT-2 tokens. Quality scores across these settings are not directly comparable. Length denotes maximum output length in each model’s tokenizer; states denotes generator sequence positions. TFLOPs reports cost per full generated document before truncation. † Fully parallel reference (Appendix C.1 ).
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.
Yulin Yuan, Ying Zhang, Xiangming Meng
Zhejiang University · University of Cambridge · ZJU-UIUC Institute, Zhejiang University
Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent diffusion modeling is constructing a suitable latent space. In this work, we present the Latent Diffusion Language Model (LDLM), in which the latent encoder, diffusion model, and decoder are trained jointly. LDLM builds its latent space by reshaping the representations of a pre-trained language model with a trainable encoder, yielding latents that are easy to both denoise and decode into tokens. We show that naive joint training produces a low-quality diffusion model, and propose a simple training recipe consisting of an MSE decoder loss, diffusion-to-encoder warmup, adaptive timestep sampling, and decoder-input noise. Ablations show that each component substantially impacts generation performance. On OpenWebText and LM1B, LDLM achieves better generation performance than existing discrete and continuous diffusion language models while being 2-13× faster, indicating that jointly learning the latent space is a key step toward making latent diffusion competitive for text generation.
Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4
Constructor University, Bremen, Germany · HSE University, Moscow, Russia · Applied AI Institute, Moscow, Russia +2
Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the same block, even though their validity depends on this realized prefix. Under the popular proposal-verification decoding, this mismatch makes errors highly asymmetric: an early rejection prevents all subsequent proposals from contributing decoding progress. We introduce BRISK-DLM, a framework that addresses both mismatches by optimizing proposal learning and selection for verified progress. BRISK-DLM trains on self-generated sequences, using risk-reward weighting to dynamically prioritize positions by their impact on verified progress and decoding cost. During inference, a lightweight prefix-conditioned corrector reranks existing candidates using previously selected tokens and preferences distilled from the model's own verifier. The corrector reuses the backbone's parallel representations and requires no additional backbone evaluation, while fused execution keeps its overhead small. BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task quality, establishing a new quality-throughput frontier for DLM generation.