Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the same block, even though their validity depends on this realized prefix. Under the popular proposal-verification decoding, this mismatch makes errors highly asymmetric: an early rejection prevents all subsequent proposals from contributing decoding progress. We introduce BRISK-DLM, a framework that addresses both mismatches by optimizing proposal learning and selection for verified progress. BRISK-DLM trains on self-generated sequences, using risk-reward weighting to dynamically prioritize positions by their impact on verified progress and decoding cost. During inference, a lightweight prefix-conditioned corrector reranks existing candidates using previously selected tokens and preferences distilled from the model's own verifier. The corrector reuses the backbone's parallel representations and requires no additional backbone evaluation, while fused execution keeps its overhead small. BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task quality, establishing a new quality-throughput frontier for DLM generation.
Figures & tables
Figure 1: From position risks to the quality–throughput frontier. Left: Parallel proposals advance generation only through the consecutively accepted prefix. Right: BRISK-DLM shifts the quality–throughput frontier toward higher throughput while preserving quality.
Figure 2
Figure 4: BRISK-DLM overview. Left: risk–reward learning and verifier distillation. Right: prefix-conditioned proposal selection followed by verification.
Figure 5: BRISK-DLM improves throughput across model sizes, tasks, and concurrency levels.
Table 5
Drafter
C=1
C=2
C=4
C=8
C=16
C=32
GM gain
Original DFlash
878.4
1612.5
2784.2
4531.7
6611.0
10046.0
–
BRISK-DFlash
925.8
2217.1
3568.7
5755.6
7141.1
10788.0
+18.3%
Table 4: Training-recipe transfer to DFlash on GSM8K. Aggregate TPS (model tokens/s) on one H200, with exactly 2,048 generated model tokens per request.
Table 7
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: First-breaker repair by proposal position. (a) Head intervention and (b) same-draw rescue, each as a fraction of factual first breakers at that position. Intervention rises from 38.4% at position 1 to 55.2% at position 7, while rescue rises from 28.2% to 35.8%. The overall upward trend suggests that later proposals offer more opportunity for prefix-conditioned correction. The rescue curve shows that the head’s additional interventions translate into more accepted local decisions, rather than merely changing token choices.
Quantity
Population / unit
Value
Candidate availability and same-state repair
Verifier top-1 in base- q Top-16
First breakers
96.97%
Corrector changes selected token
First breakers
43.07%
Corrector rescues rejection
Per 100 first breakers
28.77
Corrector introduces rejection
Per 100 accepted positions
1.60
Whole-block verifier preference
Appendix
Table 7: Mechanism diagnostics for residual selection errors. Candidate-availability and local-substitution statistics are equal-request averages over 100 GSM8K and 100 MATH-500 requests.
Figure 7: Repeat prevalence among verified positions and first breakers. Rates are equal-request averages.
Setting
8B
32B
Training data
All valid
All valid
Optimizer / LR
AdamW / 2×10−5
AdamW / 10−5
Global batch size
32
32
Sequence length
4096
8192
Training epochs
5
5
Training GPUs
8× H20
8× H20
Appendix
Table 8: Training and inference configurations.
Figure 8: Evolution of the seven normalized proposal-position weights in our experimental setting. Early positions (1–3) and later positions (4–7) are separated for clarity. Faint curves show logged values, solid curves show a centered three-point moving average, and the dashed line denotes uniform weighting ( wˉi=1 ). Emphasis initially favors early positions and gradually shifts toward later positions as the acceptance profile evolves. The panels use different vertical scales.
Setting
Value
Proposal positions / candidate size
7 / 16
Recurrent cell / state dimension
minimal GRU / 128
Hidden and token projection dimensions
64 / 64
Position / feature projection dimensions
16 / 16
Candidate residual rank
64
Trainable parameters
1.89 M
Appendix
Table 9: Candidate-refinement head configuration.
Method
MATH-500
GSM8K
HumanEval
LLaDA-2.0-flash (100B)
3.814
5.404
11.636
LLaDA-2.1-flash (100B)
4.697
6.628
13.838
SDAR-8B
1.824
1.770
1.618
SDAR-30B-A3B
1.680
1.817
1.411
DFlash OOD-thinking
3.089
3.141
5.404
DFlash native NT
6.502
7.087
1.724
Appendix
Table 11: H200 C=1 tokens per forward. A dash denotes an unavailable matched measurement.
K
C=1
C=8
C=16
4
91.70
706.19
1085.16
8
92.90
682.29
1076.39
16
94.22
741.38
1100.17
32
92.47
707.18
1012.90
Appendix
Table 14: Top- K serving throughput on A100 (TPS).
Figure 9: Cumulative reduction in head execution cost. Each bar reconstructs reference latency: two differences between successive measured paths (reference to fused scan, then to CUDA Graph replay), followed by the remaining runtime. These segments are not independent kernel timings. Labels show reference-to-optimized latency and speedup.
C
Setting
Mean prefix ↑
TPF ↑
Total TPS ↑
Request p95 (s) ↓
1
Head Off
4.20
3.37
90.0
30.20
Causal Head (Reference)
4.58
3.74
60.4
40.35
Causal Head (Optimized)
4.62
3.75
97.4
25.00
8
Head Off
4.23
3.39
682.1
24.72
Causal Head (Reference)
4.58
3.71
436.3
38.18
Causal Head (Optimized)
4.60
3.73
747.0
21.91
Appendix
Table 15: End-to-end inference-head system ablation. Absolute medians over six fresh-server runs per setting on 64 MATH-500 prompts; request p95 is computed per run. Both head implementations use the same checkpoint.