cs.LGOct 7, 2026

Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization

Authors: Xinsong Feng, Peng Du, Zhizhuo Yang, Daniel M. Bikel, Jiayun Wang, Haipeng Chen

Organizations: Georgia Institute of Technology · Writer AI Research · William & Mary

Abstract

Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emph{Branching Latent Diffusion (BLD)}, which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a 16×16\times reduction. BLD combines latent compression with \emph{branching token realization}, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than 80×80\times and increases throughput by more than 6×6\times relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than 6×6\times higher throughput and more than 4×4\times lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion

    Sep 27, 2026Yulin Yuan, Ying Zhang, Xiangming MengDiffusion Language ModelsLarge Language Model Pretraining

  2. How to Train Your Latent Diffusion Language Model Jointly With the Latent Space

    May 8, 2026Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4Diffusion Language ModelsLatent Diffusion Model

  3. From Position Risks to Block Survival: Faster Generation for Diffusion Language Models

    Sep 27, 2026Siwei Chen, Yuxiang Wan, Yifan Yu +1Diffusion Language ModelsDiffusion Models