cs.AISep 27, 2026

One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion

Authors: Yulin Yuan, Ying Zhang, Xiangming Meng

Organizations: Zhejiang University · University of Cambridge · ZJU-UIUC Institute, Zhejiang University

Abstract

Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How to Train Your Latent Diffusion Language Model Jointly With the Latent Space

    May 8, 2026Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4Diffusion Language ModelsLatent Diffusion Model

  2. Scaling and Distilling Text Embeddings for Better Diffusibility

    Oct 1, 2026Zekai Zhang, Yunjie Tian, Yanjin He +4Diffusion Language ModelsText Embeddings

  3. ELF: Embedded Language Flows

    May 11, 2026Keya Hu, Linlu Qiu, Yiyang Lu +5Diffusion Language ModelsLatent Flow