Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slower one answers correctly. On the serving engine and its decode traces, the default commit rule already commits every fully resolved block, a static skip rule captures nearly all of the compute that allocation can save, and self-distillation on engine-decoded targets adds speed at unchanged accuracy. Larger speedups come from lower thresholds, which commit tokens that are still uncertain. We therefore present Redline, a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts. Redline keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget with high probability, and deploys the fastest configuration that passes. It speeds up math at a smaller risk budget than code in both model families, and at a budget of ten percent it deploys a LLaDA2 math configuration that commits over a third more tokens in each forward. It also applies without modification to the acceptance rule of speculative decoding and to weight quantization. On the same calibration data, Redline stays within its stated failure probability, whereas each tolerance of a mean-accuracy rule either gains less speed for some model and task or exceeds the risk budget far more often for another. Code is available at https://github.com/js-lee-AI/Redline.
Figures & tables
engine accuracy
engine TPF
τadd
0.05
0.10
0.20
0.05
0.10
0.20
base
0.895
0.895
0.902
3.20
3.18
3.19
distilled
0.898
0.902
0.902
3.36
3.37
3.37
change
+0.39
+0.78
+0.00
+5.2
+5.8
+5.8
Table 1: Self-distillation on engine-decoded targets, SDAR on 256 held-out GSM8K prompts. Engine accuracy and TPF at each admission threshold τadd , with the change in points and in percent.
α=0.10
α=0.15
α=0.20
grid
n
m
αmin
gain
Δ acc.
gain
Δ acc.
gain
Δ acc.
[0pt][0pt] Block-diffusion serving, TPF gain (%)
LLaDA2 math
1,012
7
0.07
+37.5
−4.0
+37.5
−4.0
+37.5
−4.0
LLaDA2 math, large grid
1,012
50
0.07
+65.4
−3.4
+72.0
−5.9
+72.0
−5.9
LLaDA2 code
542
8
0.11
0.0
0.0
+33.7
−7.0
+42.4
−13.1
SDAR math
1,012
8
0.10
+13.1
−2.5
+36.6
−4.6
+36.6
−4.6
Table 2: Configurations that Redline deploys at three risk budgets ( δ=0.10 , Holm), with the gain over the reference in the group’s cost measure and the net accuracy change in points. Below αmin , the smallest budget at which a cheaper configuration passes, the reference is deployed.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
λ , Λ
a serving configuration and the finite grid of configurations searched, spanning accept and semi-completion thresholds, block-add schedule, buffer depth, checkpoint and precision
reference
the conservative baseline configuration of each grid, which is the engine’s default for block diffusion, lossless verification for speculative decoding and bf16 weights for quantization
R(λ)
reference-relative risk, the joint probability that the reference answers a prompt correctly and λ does not. It upper-bounds the net accuracy drop
α
the risk budget. The guarantee asserts R(λ)≤α
δ
the failure probability. All statements about one grid hold jointly with probability at least 1−δ , and δ=0.10 throughout
TPF
tokens committed in each model forward, a count ratio independent of host speed. The speculative-decoding analogue is the number of accepted tokens for each target forward
n
the number of calibration prompts of a grid
Appendix
Table 3: Notation used throughout the paper.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
[0pt][0pt] Hand-picked configurations of the SDAR math grid
acc85/semi70
5.829
147.2
858
0.106 (53/500)
1.000
0.704
0.003
1e-8
acc90/semi70
5.490
164.5
903
0.118 (59/500)
1.000
0.919
0.023
8e-7
acc85/semi90
5.390
161.2
869
0.102 (51/500)
1.000
0.596
0.001
3e-9
acc90/semi90
4.752
177.9
845
0.094 (47/500)
1.000
0.361
1e-4
9e-11
Appendix
Table 4: Redline on MATH-500 ( n=500 ) over the eight hand-picked SDAR math configurations and the distilled checkpoint at the same threshold settings ( m=16 ). TPF, forwards and tokens of each request, joint risk R^ and p -values. Bold marks configurations valid under Holm, and underlining the deployed one.
gain
net accuracy
joint risk
grid
α
value
95% interval
value
95% interval
R^
95% interval
90% bound
[0pt][0pt] Block-diffusion serving, TPF gain (%)
LLaDA2 math
≥0.10
+37.5
[34.0,41.1]
−4.0
[−5.9,−2.0]
0.072
[0.057,0.090]
0.084
LLaDA2 math, large grid
0.10
+65.4
[60.8,70.2]
−3.4
[−5.4,−1.4]
0.073
[0.058,0.091]
0.085
LLaDA2 math, large grid
≥0.15
+72.0
[67.7,76.3]
−5.9
[−8.2,−3.8]
0.098
[0.080,0.118]
0.111
LLaDA2 code
0.15
+33.7
[26.8,41.1]
−7.0
[−10.1,−3.9]
0.109
[0.084,0.138]
0.128
Appendix
Table 5: Intervals for every configuration that Redline deploys, other than the reference. Gain and net accuracy change with 95% paired-bootstrap intervals over prompts, and joint risk R^ with its 95% Clopper–Pearson interval and one-sided 90% upper bound. Budgets that deploy the same configuration share a row.
family
draw
n
median
IQR
≤0.10
<0.11
≤0.11
draws by smallest budget with a gain, 0.05 to 0.12
LLaDA2 math
uniform
542
0.08
[0.07, 0.08]
1.000
1.000
1.000
0, 52, 334, 543, 70, 1, 0, 0
LLaDA2 math
stratified
542
0.08
[0.07, 0.08]
1.000
1.000
1.000
1, 36, 324, 551, 87, 1, 0, 0
SDAR math
uniform
664
0.10
[0.10, 0.11]
0.739
0.739
0.991
0, 0, 0, 14, 224, 501, 252, 9
SDAR math
stratified
664
0.10
[0.09, 0.11]
0.741
0.741
0.994
0, 0, 0, 13, 252, 476, 253, 6
Appendix
Table 6: Smallest math budget with a gain, with the math set subsampled to the code n , over 1,000 uniform or benchmark-stratified draws for each family (Redline, δ=0.10 , fine budget grid). Median, interquartile range, shares of draws under the listed budgets, and the draws at each budget.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
acc85/semi70
6.050
128.8
779
0.072 (73/1012)
0.999
0.001
3e-14
5e-30
acc85/semi90
5.850
134.3
786
0.066 (67/1012)
0.990
1e-4
1e-16
4e-33
acc90/semi70
5.456
144.4
788
0.064 (65/1012)
0.981
4e-5
2e-17
3e-34
acc90/semi90
5.165
149.5
772
0.053 (54/1012)
0.718
6e-8
2e-22
7e-41
acc95/semi70
4.571
171.4
784
0.051 (52/1012)
0.616
1e-8
2e-23
3e-42
Appendix
Table 8: Full grid for LLaDA2 math. TPF, forwards and tokens of each request, joint risk R^ ( k/n ) and exact binomial p -values at four budgets. Bold marks the configurations valid under Holm, and underlining the one Redline deploys.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
acc85/semi70
7.566
28.9
219
0.157 (85/542)
1.000
1.000
0.697
0.006
acc85/semi90
7.107
28.6
203
0.109 (59/542)
1.000
0.778
0.003
9e-9
acc90/semi70
6.793
30.4
207
0.138 (75/542)
1.000
0.998
0.245
1e-4
acc90/semi90
6.206
34.0
211
0.076 (41/542)
0.996
0.031
1e-7
7e-16
acc95/semi70
5.495
38.3
210
0.100 (54/542)
1.000
0.525
4e-4
2e-10
Appendix
Table 9: Full grid for LLaDA2 code. TPF, forwards and tokens of each request, joint risk R^ ( k/n ) and exact binomial p -values at four budgets. Bold marks the configurations valid under Holm, and underlining the one Redline deploys.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
acc85/semi70
5.263
102.9
542
0.105 (106/1012)
1.000
0.714
2e-5
2e-16
acc90/semi70
4.956
115.0
570
0.099 (100/1012)
1.000
0.476
1e-6
2e-18
acc85/semi90
4.920
112.0
551
0.093 (94/1012)
1.000
0.244
4e-8
1e-20
acc90/semi90
4.361
123.9
540
0.071 (72/1012)
0.999
9e-4
1e-14
2e-30
acc95/semi70
4.296
136.0
584
0.083 (84/1012)
1.000
0.037
9e-11
8e-25
Appendix
Table 10: Full grid for SDAR math. TPF, forwards and tokens of each request, joint risk R^ ( k/n ) and exact binomial p -values at four budgets. Bold marks the configurations valid under Holm, and underlining the one Redline deploys.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
acc85/semi70
4.253
17.2
73
0.178 (118/664)
1.000
1.000
0.978
0.081
acc85/semi90
3.810
18.0
69
0.145 (96/664)
1.000
1.000
0.372
1e-4
acc90/semi70
3.492
19.4
68
0.139 (92/664)
1.000
0.999
0.222
2e-5
acc90/semi90
3.349
20.6
69
0.087 (58/664)
1.000
0.153
9e-7
1e-15
acc95/semi70
2.980
23.7
71
0.081 (54/664)
1.000
0.059
7e-8
3e-17
Appendix
Table 11: Full grid for SDAR code. TPF, forwards and tokens of each request, joint risk R^ ( k/n ) and exact binomial p -values at four budgets. Bold marks the configurations valid under Holm, and underlining the one Redline deploys.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
acc85/semi70
5.973
131.7
787
0.057 (29/512)
0.789
3e-4
3e-11
2e-20
acc85/semi90
5.850
129.8
759
0.070 (36/512)
0.983
0.012
2e-8
2e-16
acc90/semi70
5.317
146.2
778
0.064 (33/512)
0.941
0.003
2e-9
5e-18
acc90/semi90
5.245
150.9
792
0.053 (27/512)
0.659
8e-5
3e-12
1e-21
acc95/semi70
4.534
165.4
750
0.051 (26/512)
0.584
4e-5
1e-12
2e-22
Appendix
Table 12: Full grid for LLaDA2 math on the reduced prompt set. TPF, forwards and tokens of each request, joint risk R^ ( k/n ) and exact binomial p -values at four budgets. Bold marks the configurations valid under Holm, and underlining the one Redline deploys.
p -value at budget α
configuration
TPF
forwards
tokens
R^ ( k/n )
0.05
0.10
0.15
0.20
acc85/semi70
7.619
28.6
218
0.143 (60/420)
1.000
0.998
0.372
0.001
acc85/semi90
7.224
31.8
230
0.121 (51/420)
1.000
0.936
0.055
1e-5
acc90/semi70
6.536
32.7
214
0.117 (49/420)
1.000
0.887
0.030
4e-6
acc90/semi90
6.339
34.2
217
0.090 (38/420)
1.000
0.290
2e-4
7e-10
acc95/semi70
5.492
40.1
220
0.093 (39/420)
1.000
0.349
3e-4
2e-9
Appendix
Table 13: Full grid for LLaDA2 code on the reduced prompt set. TPF, forwards and tokens of each request, joint risk R^ ( k/n ) and exact binomial p -values at four budgets. Bold marks the configurations valid under Holm, and underlining the one Redline deploys.
Redline
MEAN ( t ) just slower
MEAN ( t ) at or faster
MEAN ( 10 )
α
gain
exc.
ref.
t
gain
exc.
t
gain
exc.
gain
exc.
LLaDA2 math
0.05
0.0
0.1
99.9
none slower
0
0.6
5.7
37.5
99.6
0.10
36.6
0.1
0.0
5
36.6
0.1
6
37.3
0.1
37.5
0.1
0.15
37.5
0.0
0.0
6
37.3
0.0
7
37.5
0.0
37.5
0.0
0.20
37.5
0.0
0.0
6
37.3
0.0
7
37.5
0.0
37.5
0.0
Appendix
Table 14: Redline against the mean-accuracy rule MEAN( t ) on 1,000 random half splits. Mean TPF gain, exceedance (splits with test-half R>α ) and deployments of the reference (ref.), all in percent, for the tolerances that bracket the gain of Redline and for t=10 .
NETHOLM
MEAN ( t )
α
measure
Redline
BONF
FIXSEQ
UNCORR
PLUGIN
H
HB
WSR
MEANCI
slower
faster
t=10
[0pt][0pt] LLaDA2 math ( n=1,012 )
0.05
gain
0.0
0.0
0.1
0.2
7.8
0.0
0.0
8.6
27.0
none
0.6
37.5
test
0.1
0.1
1.6
2.4
58.4
0.0
0.0
47.9
89.6
5.7
99.6
pool
0.1
0.1
1.6
2.4
58.4
0.0
0.0
50.3
99.5
5.7
100.0
0.10
gain
36.6
32.7
36.7
37.0
37.5
0.0
0.0
37.5
37.5
36.6
37.3
37.5
Appendix
Table 15: Every selector on the same 1,000 random half splits ( δ=0.10 ). Mean TPF gain of the deployed configuration, and exceedance on the test half and against the pooled risk, in percent.
in sample
splits
held out
α
deployed
gain
net
ref.
same
gain
net
LLaDA2 math
0.05
reference
+0.0
0.0
999
999
+0.0
0.0
0.10
acc85/semi70
+37.5
− 4.0
0
894
+37.6 [+34.1, +41.3]
− 4.2 [ − 5.9, − 2.6]
0.15
acc85/semi70
+37.5
− 4.0
0
1000
+37.6 [+34.1, +41.4]
− 4.0 [ − 5.9, − 2.2]
0.20
acc85/semi70
+37.5
− 4.0
0
1000
+37.6 [+34.1, +41.4]
− 4.0 [ − 5.9, − 2.2]
Appendix
Table 16: Held-out re-measurement over 1,000 random half splits of the configuration that Redline deploys on one half, measured on the other. Splits deploying the reference and the full-sample configuration, and held-out mean with 95% percentile range, gain in percent and net accuracy in points.
α
deployed
valid
k
R^
TPF
TPF gain (%)
net accuracy (pp)
[0pt][0pt] The six configurations shared with the main grid, m=6
0.05
reference
1 of 6
0
0.000
4.42
+0.0
+0.0
0.10
acc85/semi70/add0.10
6 of 6
53
0.052
5.90
+33.4
−1.5
0.15
acc85/semi70/add0.10
6 of 6
53
0.052
5.90
+33.4
−1.5
0.20
acc85/semi70/add0.10
6 of 6
53
0.052
5.90
+33.4
−1.5
[0pt][0pt] The add 0.10 slice, m=25
Appendix
Table 17: The LLaDA2 math grid of 50 configurations tested by Redline as three nested grids of the same served outputs ( n=1,012 , δ=0.10 ). For each grid and budget, the deployed configuration, the number of valid configurations, its violations k , joint risk, TPF, gain and net accuracy change.
line of work
guarantee
multiplicity
mechanism
generality
BD-LM families and serving engine (BD3-LM, SDAR, LLaDA2.0 and 2.1, MBD-LM)
✗mean accuracy
✗hand-picked points
stall documented, unexplained
✗
learned-frontier line (distillation, multi-block visibility, RL shaping, in-place revision, Fast-dLLM and v2, D2F, AdaBlock, dParallel, LoPA, d3LLM, DMax)
✗
✗
✗
one cost each
Fast-dLLM Theorem 1
deterministic, conditional on trusting model confidence, no δ
✗
✗
one threshold
CALM (LTT on AR early exit)
✓finite-sample, relative to the full model
one threshold
✗
early exit only
LTT and conformal risk control (frameworks)
✓
✓procedure
not a serving system
lossy AR acceleration (speculative decoding, Medusa typical acceptance, PTQ)
✗supply the settings
✗
✗
one cost each
Appendix
Table 18: Positioning against related work. Guarantee is a distribution-free finite-sample statement about deployed quality, multiplicity is error control over a space of configurations, mechanism is causal evidence for the operating point, and generality is one procedure across cost measures.
Speculative decoding speeds up LLM inference by using a draft model to generate tokens, with an acceptance-rejection scheme that ensures that the output matches the target distribution. Adapting this to continuous diffusions is difficult because speculative sampling requires drawing from a residual distribution. While straightforward in discrete spaces, efficiently sampling this residual in continuous space is non-trivial. Consequently, existing diffusion adaptations either use computationally inefficient sampling techniques or rely on an alternative scheme. In this work, we introduce a novel scheme that efficiently implements the original speculative sampling mechanism for diffusion models. Our approach offers a critical advantage over current methods: it enables us to adapt block verification from LLMs to diffusions -- which provably improves the acceptance rate of drafts. Furthermore, we formalize and analyze the Free Drafter, a heuristic self-speculative drafter for diffusions that requires no training. By enabling block verification, our Free Drafter yields up to a 6.3% speedup over existing speculative methods with no additional training and negligible overhead beyond the existing parallel verification pass.
Alexander Soen, Hisham Husain, Valentin De Bortoli +1
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion. Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations. In Flash-Flash, the same Flash model serves as both drafter and refiner, testing whether an existing model can improve its own block-autoregressive output through global refinement. In Mini-Flash, inspired by speculative decoding, we introduce speculative correction: Mini drafts a full response, and Flash revises it as an editable initialization. Flash-Flash improves GSM8K-384 accuracy from 0.848 to 0.899 while running 1.20 times faster than the selected Flash block-autoregressive baseline, and improves MBPP-384 from 0.545 to 0.693. Latency-window-matched Flash-only controls indicate that these gains persist after targeted tuning of block-autoregressive decoding. Causal ablations indicate that completed drafts provide useful initializations: refinement from a fully masked span performs poorly, full global refinement provides a clear additional gain on GSM8K, and local refinement captures much of the gain on MBPP and MATH. Mini-Flash provides useful quality-latency trade-offs, including MATH-384 performance of 0.294 versus 0.300 for Flash while running 2.17 times faster. These results support a Pareto-frontier interpretation rather than the claim that the heterogeneous cascade uniformly matches Flash quality. Overall, same-model draft-and-refine provides evidence that bidirectional refinement is a useful decoding primitive for DLMs, while speculative correction demonstrates a training-free route to fast DLM generation.
Brian K Chen, Chong Wu, Kenji Kawaguchi
National University of Singapore · City University of Hong Kong
Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabilizes before the step schedule is exhausted. This creates two acceleration opportunities, leaving a block early and stopping the sequence early, but the two require different criteria because block acceleration is local whereas sequence termination is global and freezes the graded answer. Existing methods usually optimize only one axis, and existing exit gates rely on fixed-region confidence or schedule-dependent rules rather than the candidate answer itself. We present C4, which coordinates the two axes by giving each decision its own gate. Confidence-Verified Early Exit (CVEE) decides when the sequence may stop, requiring confidence and sustained argmax stability over a candidate span re-extracted at every step. Commit-Core-Then-Confirm (CCTC) decides which token positions a step may commit by borrowing an autoregressive freezing order inside each block: it commits a boundary-anchored core and confirms deferred positions one step later, so the answer block can be accelerated without allowing local commits inside the answer span to determine sequence-level termination. On 12 zero-shot tasks with LLaDA and Dream, one frozen configuration removes 64--95% of decoding steps and delivers measured end-to-end speedups of 2.6 to 8.6 over full decoding. Code is available at https://github.com/ming053l/C4-dLLM.
Chia-Ming Lee, Shao-Kai Liu, Ming-Ching Chang +3
National Yang Ming Chiao Tung University · University at Albany, SUNY