Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the same block, even though their validity depends on this realized prefix. Under the popular proposal-verification decoding, this mismatch makes errors highly asymmetric: an early rejection prevents all subsequent proposals from contributing decoding progress. We introduce BRISK-DLM, a framework that addresses both mismatches by optimizing proposal learning and selection for verified progress. BRISK-DLM trains on self-generated sequences, using risk-reward weighting to dynamically prioritize positions by their impact on verified progress and decoding cost. During inference, a lightweight prefix-conditioned corrector reranks existing candidates using previously selected tokens and preferences distilled from the model's own verifier. The corrector reuses the backbone's parallel representations and requires no additional backbone evaluation, while fused execution keeps its overhead small. BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task quality, establishing a new quality-throughput frontier for DLM generation.
Figures & tables
Figure 1: From position risks to the quality–throughput frontier. Left: Parallel proposals advance generation only through the consecutively accepted prefix. Right: BRISK-DLM shifts the quality–throughput frontier toward higher throughput while preserving quality.
Figure 2
Figure 4: BRISK-DLM overview. Left: risk–reward learning and verifier distillation. Right: prefix-conditioned proposal selection followed by verification.
Figure 5: BRISK-DLM improves throughput across model sizes, tasks, and concurrency levels.
Table 5
Drafter
C=1
C=2
C=4
C=8
C=16
C=32
GM gain
Original DFlash
878.4
1612.5
2784.2
4531.7
6611.0
10046.0
–
BRISK-DFlash
925.8
2217.1
3568.7
5755.6
7141.1
10788.0
+18.3%
Table 4: Training-recipe transfer to DFlash on GSM8K. Aggregate TPS (model tokens/s) on one H200, with exactly 2,048 generated model tokens per request.
Table 7
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: First-breaker repair by proposal position. (a) Head intervention and (b) same-draw rescue, each as a fraction of factual first breakers at that position. Intervention rises from 38.4% at position 1 to 55.2% at position 7, while rescue rises from 28.2% to 35.8%. The overall upward trend suggests that later proposals offer more opportunity for prefix-conditioned correction. The rescue curve shows that the head’s additional interventions translate into more accepted local decisions, rather than merely changing token choices.
Quantity
Population / unit
Value
Candidate availability and same-state repair
Verifier top-1 in base- q Top-16
First breakers
96.97%
Corrector changes selected token
First breakers
43.07%
Corrector rescues rejection
Per 100 first breakers
28.77
Corrector introduces rejection
Per 100 accepted positions
1.60
Whole-block verifier preference
Appendix
Table 7: Mechanism diagnostics for residual selection errors. Candidate-availability and local-substitution statistics are equal-request averages over 100 GSM8K and 100 MATH-500 requests.
Figure 7: Repeat prevalence among verified positions and first breakers. Rates are equal-request averages.
Setting
8B
32B
Training data
All valid
All valid
Optimizer / LR
AdamW / 2×10−5
AdamW / 10−5
Global batch size
32
32
Sequence length
4096
8192
Training epochs
5
5
Training GPUs
8× H20
8× H20
Appendix
Table 8: Training and inference configurations.
Figure 8: Evolution of the seven normalized proposal-position weights in our experimental setting. Early positions (1–3) and later positions (4–7) are separated for clarity. Faint curves show logged values, solid curves show a centered three-point moving average, and the dashed line denotes uniform weighting ( wˉi=1 ). Emphasis initially favors early positions and gradually shifts toward later positions as the acceptance profile evolves. The panels use different vertical scales.
Setting
Value
Proposal positions / candidate size
7 / 16
Recurrent cell / state dimension
minimal GRU / 128
Hidden and token projection dimensions
64 / 64
Position / feature projection dimensions
16 / 16
Candidate residual rank
64
Trainable parameters
1.89 M
Appendix
Table 9: Candidate-refinement head configuration.
Method
MATH-500
GSM8K
HumanEval
LLaDA-2.0-flash (100B)
3.814
5.404
11.636
LLaDA-2.1-flash (100B)
4.697
6.628
13.838
SDAR-8B
1.824
1.770
1.618
SDAR-30B-A3B
1.680
1.817
1.411
DFlash OOD-thinking
3.089
3.141
5.404
DFlash native NT
6.502
7.087
1.724
Appendix
Table 11: H200 C=1 tokens per forward. A dash denotes an unavailable matched measurement.
K
C=1
C=8
C=16
4
91.70
706.19
1085.16
8
92.90
682.29
1076.39
16
94.22
741.38
1100.17
32
92.47
707.18
1012.90
Appendix
Table 14: Top- K serving throughput on A100 (TPS).
Figure 9: Cumulative reduction in head execution cost. Each bar reconstructs reference latency: two differences between successive measured paths (reference to fused scan, then to CUDA Graph replay), followed by the remaining runtime. These segments are not independent kernel timings. Labels show reference-to-optimized latency and speedup.
C
Setting
Mean prefix ↑
TPF ↑
Total TPS ↑
Request p95 (s) ↓
1
Head Off
4.20
3.37
90.0
30.20
Causal Head (Reference)
4.58
3.74
60.4
40.35
Causal Head (Optimized)
4.62
3.75
97.4
25.00
8
Head Off
4.23
3.39
682.1
24.72
Causal Head (Reference)
4.58
3.71
436.3
38.18
Causal Head (Optimized)
4.60
3.73
747.0
21.91
Appendix
Table 15: End-to-end inference-head system ablation. Absolute medians over six fresh-server runs per setting on 64 MATH-500 prompts; request p95 is computed per run. Both head implementations use the same checkpoint.
Diffusion language models (DLMs) offer substantial speed advantages through parallel decoding, but the lack of token dependencies limits generation quality compared to autoregressive (AR) models. Recent progress attempts to bridge the gap via importance sampling, with DLM being the proposal and AR being the target. However, due to the huge gap between their distributions, the sampling requires a large number of particles and is thus expensive to compute. In this paper, we introduce PoE-Bridge, a novel decoding framework that drastically improves generation speed and accuracy by introducing an intermediate distribution to bridge the gap. The distribution is constructed as a Product-of-Experts (PoE) of the DLM proposal and the AR target. With the intermediate distribution, we first use the DLM to draft multiple continuations in parallel, then apply rejection sampling to verify the drafted tokens and move the resulting candidates toward the PoE. We then use importance sampling to further correct the PoE-aligned candidates toward the AR target. We further propose several improved techniques, including mixed-temperature sampling for enhanced diversity and elastic rejection windows for reducing wasted verification. Empirically, PoE-Bridge achieves significantly improved accuracy with 5× speedup over the standard DLM decoding approach, and recovers at least 95% of the target AR model's performance, efficiently advancing most of the quality gap on challenging mathematical reasoning and coding tasks. Our code is available at https://github.com/juntongshi48/poe-bridge.
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion. Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations. In Flash-Flash, the same Flash model serves as both drafter and refiner, testing whether an existing model can improve its own block-autoregressive output through global refinement. In Mini-Flash, inspired by speculative decoding, we introduce speculative correction: Mini drafts a full response, and Flash revises it as an editable initialization. Flash-Flash improves GSM8K-384 accuracy from 0.848 to 0.899 while running 1.20 times faster than the selected Flash block-autoregressive baseline, and improves MBPP-384 from 0.545 to 0.693. Latency-window-matched Flash-only controls indicate that these gains persist after targeted tuning of block-autoregressive decoding. Causal ablations indicate that completed drafts provide useful initializations: refinement from a fully masked span performs poorly, full global refinement provides a clear additional gain on GSM8K, and local refinement captures much of the gain on MBPP and MATH. Mini-Flash provides useful quality-latency trade-offs, including MATH-384 performance of 0.294 versus 0.300 for Flash while running 2.17 times faster. These results support a Pareto-frontier interpretation rather than the claim that the heterogeneous cascade uniformly matches Flash quality. Overall, same-model draft-and-refine provides evidence that bidirectional refinement is a useful decoding primitive for DLMs, while speculative correction demonstrates a training-free route to fast DLM generation.
Brian K Chen, Chong Wu, Kenji Kawaguchi
National University of Singapore · City University of Hong Kong
Diffusion Language Models (DLMs) generate text by iteratively denoising masked token sequences, offering a tradeoff between parallelism and quality compared to autoregressive models. In current practice, the number of tokens decoded per step is controlled by a confidence threshold, and quality degrades monotonically as more tokens are denoised per step. We introduce Multi-token Residual Prediction (MRP), a lightweight module that enables dependency-aware multi-token denoising within a single backbone forward pass. MRP exploits a key property of the denoising process: the logit distributions at adjacent denoising steps are remarkably similar. Rather than running the backbone a second time to obtain the next-step logits, MRP predicts the residual between steps from the backbone's hidden states, effectively denoising more tokens per backbone forward at a fraction of the cost. We apply MRP across the two operating regimes of DLM decoding. In the high-quality-low-throughput static denoising regime, MRP serves as a drafter for speculative decoding: its proposals are verified against the backbone, yielding lossless acceleration of up to 1.4x in SGLang. In the low-quality-high-throughput dynamic denoising regime, MRP instead drives a remasking scheme that revokes over-eager reveals, recovering most of the accuracy lost to aggressive low-threshold decoding and improving accuracy by up to 22.6 points on code generation task HumanEval and 17.7 points on reasoning task GSM8K.
Yufeng Xu, Zishuo Bao, Qian Wang +6
New York University · New York University Shanghai · Nous Research +1