Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.
Figures & tables
Figure 1: Schematic comparison of fixed and jointly learned compressed spaces, showing more structured embeddings and broader decodable regions with joint training.
Figure 2: Overview of JPEG-DLM . The frozen T5 encoder maps tokens to features. The context compressor maps a corrupted view to zc , while the target compressor provides the stop-gradient clean target embedding ztgt and is updated by an exponential moving average. The shared-weight network predicts the clean target embedding from zc for JEP and from the noisy embedding zt for flow matching. The decoding module maps the predicted clean embedding back to features and tokens using Lfeat and Lce .
Figure 3
LM1B ( L=128 )
OWT ( L=1024 )
Method
Param. (M)
Gen-PPL ↓
Ent.
MAUVE ↑
Speed ↑
Gen-PPL ↓
Ent.
MAUVE ↑
Speed ↑
AR ( Radford et al., 2019 )
130
151.36
4.397
0.947
1.00× ( 1.62× w/ KV cache)
41.06
5.583
0.930
1.00× ( 7.14× w/ KV cache)
MDLM ( Sahoo et al., 2024 )
130
192.80
4.410
0.923
0.52×
41.46
5.287
0.755
0.32×
SEDD-Absorb ( Lou et al., 2024 )
170
182.57
4.405
0.941
0.40×
41.15
5.260
0.728
0.25×
Duo ( Sahoo et al., 2025 )
130
161.80
4.382
0.941
0.41×
75.88
5.536
0.903
0.15×
FLM ( Lee et al., 2026 )
179
176.74
4.402
0.863
0.66×
61.37
5.335
0.373
0.21×
Table 1: Unconditional generation on LM1B and OWT. Speed is relative to AR without KV cache. Among diffusion and flow models, bold and underlined mark the best and second-best Gen-PPL, MAUVE, and speed. JPEG-DLM has 105M DiT and token projection parameters, or 154M including the compressor and decompressor.
Figure 3: Generation quality at different latent lengths (a) and latent widths (b).
Figure 4: Joint and two-stage training. (a) Gen-PPL and MAUVE across compression rates; arrows show relative changes from COSMOS to JPEG-DLM , and diamonds mark ELF-B. COSMOS uses final-layer T5-small features at every rate. (b) Correlation between latent positions at r=0.5 .
Figure 5: Gen-PPL over interpolations between 1,024 pairs of clean target embeddings at a compression rate of r=0.5 , without added noise and at three noise levels. Here, t denotes the COSMOS forward time, with t=0 indicating no added noise. For the three noisy settings, JPEG-DLM uses noisy embeddings at the same SNRs along its linear path.
Variant
Ljep
Lfeat
σ
Lce timesteps
Gen-PPL ↓
Ent.
MAUVE ↑
Base
✓
✓
1.0
t=1
48.00
4.064
0.912
sep. nets
✓
✓
1.0
t=1
92.45
4.117
0.786
w/o Ljep
×
✓
1.0
t=1
51.65
4.016
0.849
w/o Ljep + sep. nets
×
✓
1.0
t=1
131.94
4.188
0.632
w/o Lfeat
✓
×
1.0
t=1
51.57
4.069
0.825
σ=0.5
✓
✓
0.5
t=1
61.43
3.993
0.817
Table 2: Ablation study on OWT-128. Base uses a shared-weight network.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
SIGReg
VICReg
Gen-PPL ↓
MAUVE ↑
Entropy
×
×
63.11
0.856
4.090
✓
×
48.60
0.865
4.062
×
✓
40.77
0.868
3.965
✓
✓
48.00
0.912
4.064
Appendix
Table 3: Regularizer ablation on OWT-128 with r=0.5 , d=512 , 64 sampling steps, and self-conditioning at scale 3.0 . The variance and covariance terms are enabled or disabled as a pair. The shaded row reproduces Base from Table 2 .
Figure 6: Gen-PPL on LM1B across sampling steps. Error bars show standard deviations over five seeds. Standard deviations are unavailable for the 8-step baselines.
Steps
Gen-PPL ↓
Ent.
MAUVE ↑
8
225.24
4.162
0.432
16
137.86
4.270
0.896
32
96.76
4.263
0.951
64
78.64
4.248
0.940
128
69.45
4.238
0.941
256
64.07
4.230
0.933
Appendix
Table 4: Generation quality of JPEG-DLM on LM1B with r=0.25 at different sampling budgets. Results are means over five seeds. The shaded row is the 32-step configuration used in the main comparison.
Figure 7: GPT-2 token lengths of ELF generations and two OWT reference sets. The reference constructed from fixed 1,024-token blocks concentrates near the block length, whereas the reference tokenized and decoded with T5 and the ELF samples vary in length. After discretization, the intersection between the ELF and fixed-length reference histograms is only 0.9%.
Parameter
LM1B
OWT-1024
Data and tokenization
Dataset
LM1B ( Chelba et al., 2014 )
OWT
Token sequence length
128
1024
Sequence construction
Blocks packed across sentence boundaries
Blocks packed across sentence boundaries (no padding)
Model architecture
Frozen text encoder
T5-small
T5-small
Appendix
Table 5: Model and training configurations of JPEG-DLM on LM1B and OWT-1024. Sampling compute counts one multiply–accumulate as one operation. For K solver steps, JPEG-DLM uses K+1 DiT evaluations, including the final clean embedding prediction, followed by one decoding pass. The estimate at full latent length uses the same K+1 evaluation budget.
LM1B
OWT-1024
Method
SC
Samples
Tokenizer
Steps
Tokenizer
Steps
AR
×
1,024
BERT
–
GPT-2
–
MDLM
×
1,024
BERT
128
GPT-2 + mask
1024
SEDD-Absorb
×
1,024
GPT-2
128
GPT-2
1024
Duo
×
1,024
BERT
128
GPT-2
1024
FLM
×
1,024
BERT
128
GPT-2
1024
Appendix
Table 6: Tokenization and sampling configurations for the main unconditional generation comparison. BERT ( Devlin et al., 2019 ) denotes bert-base-uncased , and GPT-2 ( Radford et al., 2019 ) denotes GPT-2 BPE. T5 ( Raffel et al., 2020 ) uses SentencePiece ( Kudo and Richardson, 2018 ) . A checkmark in the SC column indicates self-conditioning.
Parameter
OWT-128
Dataset
linluqiu/openwebtext-len128-packing_padding
Sequence length
128
Text encoder / tokenizer
Frozen T5-small / T5
Encoder feature
Final hidden layer
Latent length / compression rate r
128 / 1.0, 64 / 0.5, 32 / 0.25, 16 / 0.125
Latent width
512
Appendix
Table 7: Configuration of COSMOS on OWT-128.
Layer index
Gen-PPL ↓
Ent.
MAUVE ↑
−1
240.811
3.9033
0.1167
−2
154.551
3.9736
0.3009
−3
163.196
3.9717
0.2677
Appendix
Table 8: COSMOS performance at r=1.0 with T5-small layer indices −1 , −2 , and −3 . Results are means over five seeds using 1,024 samples per seed.
Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent diffusion modeling is constructing a suitable latent space. In this work, we present the Latent Diffusion Language Model (LDLM), in which the latent encoder, diffusion model, and decoder are trained jointly. LDLM builds its latent space by reshaping the representations of a pre-trained language model with a trainable encoder, yielding latents that are easy to both denoise and decode into tokens. We show that naive joint training produces a low-quality diffusion model, and propose a simple training recipe consisting of an MSE decoder loss, diffusion-to-encoder warmup, adaptive timestep sampling, and decoder-input noise. Ablations show that each component substantially impacts generation performance. On OpenWebText and LM1B, LDLM achieves better generation performance than existing discrete and continuous diffusion language models while being 2-13× faster, indicating that jointly learning the latent space is a key step toward making latent diffusion competitive for text generation.
Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4
Constructor University, Bremen, Germany · HSE University, Moscow, Russia · Applied AI Institute, Moscow, Russia +2
Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today's leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs can be made effective with minimal adaptation to the discrete domain. We propose Embedded Language Flows (ELF), a class of diffusion models in continuous embedding space based on continuous-time Flow Matching. Unlike existing DLMs, ELF predominantly stays within the continuous embedding space until the final time step, where it maps to discrete tokens using a shared-weight network. This formulation makes it straightforward to adapt established techniques from image-domain diffusion models, e.g., classifier-free guidance (CFG). Experiments show that ELF substantially outperforms leading discrete and continuous DLMs, achieving better generation quality with fewer sampling steps. These results suggest that ELF offers a promising path toward effective continuous DLMs.