cs.CLOct 5, 2026

D-Loop: Looped Diffusion Drafting for Speculative Decoding

Authors: Kecheng Chen, Yuyang He, Cheng Gong, Hui Liu, Guoping Long, Jiajun Li, Shi Wu, Suiyun Zhang, +3 more

Organizations: City University of Hong Kong · The Chinese University of Hong Kong · Huawei Research

Abstract

Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.

Figures & tables

Explore similar work

CardsList
  1. xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

    Aug 3, 2026Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4Block-Diffusion DrafterDrafts

  2. D^2SD: Accelerating Speculative Decoding with Dual Diffusion Draft Models

    Jun 3, 2026Liyuan Zhang, Jiarui Zhang, Jinwei Yao +6Speculative DecodingBlock-Diffusion Drafter

  3. DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding

    Jun 1, 2026Jiebin Zhang, Zhenghan Yu, Song Liu +9Block-Diffusion DrafterSpeculative Decoding