cs.CLOct 5, 2026

Representation-Space MMD for Diffusion Language Models

Authors: Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev, Maksim Ignatov, Pavel Temirchev, Nikita Balagansky, Viacheslav Meshchaninov, +2 more

Organizations: Yandex Research · HSE University · T-Tech · Constructor University · Applied AI Institute · AXXX

Abstract

We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How to Train Your Latent Diffusion Language Model Jointly With the Latent Space

    May 8, 2026Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4Diffusion Language ModelsLatent Diffusion Model

  2. Distribution Matching Distillation for Continuous Diffusion Language Models

    Sep 30, 2026Paul Le Van Kiem, Dario Shariatian, Umut Simsekli +1Diffusion Language ModelsDiffusion Models