cs.LGSep 30, 2026

Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models

Authors: Stipe Frković, Metod Jazbec, Christian A. Naesseth

Organizations: University of Amsterdam · UvA-Bosch Delta Lab, University of Amsterdam

Abstract

Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Parallelism, critical windows, and separations among diffusion language models

    Sep 17, 2026Sitan Chen, Liye WangDiffusion Language ModelsDiffusion Sampling

  2. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

    Jun 9, 2026Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher +1Masked Diffusion Language ModelsDiffusion Language Models

  3. Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models

    Jun 10, 2026Jia Deng, Junyi Li, Wayne Xin Zhao +3Diffusion Language Models