Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees
Organizations: Korea University · Zoom Communications · Soongsil University · Yonsei University Mirae Campus
Abstract
Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slower one answers correctly. On the serving engine and its decode traces, the default commit rule already commits every fully resolved block, a static skip rule captures nearly all of the compute that allocation can save, and self-distillation on engine-decoded targets adds speed at unchanged accuracy. Larger speedups come from lower thresholds, which commit tokens that are still uncertain. We therefore present Redline, a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts. Redline keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget with high probability, and deploys the fastest configuration that passes. It speeds up math at a smaller risk budget than code in both model families, and at a budget of ten percent it deploys a LLaDA2 math configuration that commits over a third more tokens in each forward. It also applies without modification to the acceptance rule of speculative decoding and to weight quantization. On the same calibration data, Redline stays within its stated failure probability, whereas each tolerance of a mean-accuracy rule either gains less speed for some model and task or exceeds the risk budget far more often for another. Code is available at https://github.com/js-lee-AI/Redline.
Figures & tables
| engine accuracy | engine TPF | |||||
| 0.05 | 0.10 | 0.20 | 0.05 | 0.10 | 0.20 | |
| base | 0.895 | 0.895 | 0.902 | 3.20 | 3.18 | 3.19 |
| distilled | 0.898 | 0.902 | 0.902 | 3.36 | 3.37 | 3.37 |
| change | ||||||
| grid | gain | acc. | gain | acc. | gain | acc. | |||
| [0pt][0pt] Block-diffusion serving, TPF gain (%) | |||||||||
| LLaDA2 math | |||||||||
| LLaDA2 math, large grid | |||||||||
| LLaDA2 code | |||||||||
| SDAR math | |||||||||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| , | a serving configuration and the finite grid of configurations searched, spanning accept and semi-completion thresholds, block-add schedule, buffer depth, checkpoint and precision |
| reference | the conservative baseline configuration of each grid, which is the engine’s default for block diffusion, lossless verification for speculative decoding and bf16 weights for quantization |
| reference-relative risk, the joint probability that the reference answers a prompt correctly and does not. It upper-bounds the net accuracy drop | |
| the risk budget. The guarantee asserts | |
| the failure probability. All statements about one grid hold jointly with probability at least , and throughout | |
| TPF | tokens committed in each model forward, a count ratio independent of host speed. The speculative-decoding analogue is the number of accepted tokens for each target forward |
| the number of calibration prompts of a grid |
| -value at budget | ||||||||
| configuration | TPF | forwards | tokens | ( ) | ||||
| [0pt][0pt] Hand-picked configurations of the SDAR math grid | ||||||||
| acc85/semi70 | 5.829 | 147.2 | 858 | 0.106 (53/500) | 1.000 | 0.704 | 0.003 | 1e-8 |
| acc90/semi70 | 5.490 | 164.5 | 903 | 0.118 (59/500) | 1.000 | 0.919 | 0.023 | 8e-7 |
| acc85/semi90 | 5.390 | 161.2 | 869 | 0.102 (51/500) | 1.000 | 0.596 | 0.001 | 3e-9 |
| acc90/semi90 | 4.752 | 177.9 | 845 | 0.094 (47/500) | 1.000 | 0.361 | 1e-4 | 9e-11 |
| gain | net accuracy | joint risk | ||||||
| grid | value | interval | value | interval | interval | bound | ||
| [0pt][0pt] Block-diffusion serving, TPF gain (%) | ||||||||
| LLaDA2 math | ||||||||
| LLaDA2 math, large grid | ||||||||
| LLaDA2 math, large grid | ||||||||
| LLaDA2 code | ||||||||
| family | draw | median | IQR | draws by smallest budget with a gain, to | ||||
|---|---|---|---|---|---|---|---|---|
| LLaDA2 math | uniform | 542 | 0.08 | [0.07, 0.08] | 1.000 | 1.000 | 1.000 | 0, 52, 334, 543, 70, 1, 0, 0 |
| LLaDA2 math | stratified | 542 | 0.08 | [0.07, 0.08] | 1.000 | 1.000 | 1.000 | 1, 36, 324, 551, 87, 1, 0, 0 |
| SDAR math | uniform | 664 | 0.10 | [0.10, 0.11] | 0.739 | 0.739 | 0.991 | 0, 0, 0, 14, 224, 501, 252, 9 |
| SDAR math | stratified | 664 | 0.10 | [0.09, 0.11] | 0.741 | 0.741 | 0.994 | 0, 0, 0, 13, 252, 476, 253, 6 |
| -value at budget | ||||||||
|---|---|---|---|---|---|---|---|---|
| configuration | TPF | forwards | tokens | ( ) | ||||
| acc85/semi70 | 6.050 | 128.8 | 779 | 0.072 (73/1012) | 0.999 | 0.001 | 3e-14 | 5e-30 |
| acc85/semi90 | 5.850 | 134.3 | 786 | 0.066 (67/1012) | 0.990 | 1e-4 | 1e-16 | 4e-33 |
| acc90/semi70 | 5.456 | 144.4 | 788 | 0.064 (65/1012) | 0.981 | 4e-5 | 2e-17 | 3e-34 |
| acc90/semi90 | 5.165 | 149.5 | 772 | 0.053 (54/1012) | 0.718 | 6e-8 | 2e-22 | 7e-41 |
| acc95/semi70 | 4.571 | 171.4 | 784 | 0.051 (52/1012) | 0.616 | 1e-8 | 2e-23 | 3e-42 |
| -value at budget | ||||||||
|---|---|---|---|---|---|---|---|---|
| configuration | TPF | forwards | tokens | ( ) | ||||
| acc85/semi70 | 7.566 | 28.9 | 219 | 0.157 (85/542) | 1.000 | 1.000 | 0.697 | 0.006 |
| acc85/semi90 | 7.107 | 28.6 | 203 | 0.109 (59/542) | 1.000 | 0.778 | 0.003 | 9e-9 |
| acc90/semi70 | 6.793 | 30.4 | 207 | 0.138 (75/542) | 1.000 | 0.998 | 0.245 | 1e-4 |
| acc90/semi90 | 6.206 | 34.0 | 211 | 0.076 (41/542) | 0.996 | 0.031 | 1e-7 | 7e-16 |
| acc95/semi70 | 5.495 | 38.3 | 210 | 0.100 (54/542) | 1.000 | 0.525 | 4e-4 | 2e-10 |
| -value at budget | ||||||||
|---|---|---|---|---|---|---|---|---|
| configuration | TPF | forwards | tokens | ( ) | ||||
| acc85/semi70 | 5.263 | 102.9 | 542 | 0.105 (106/1012) | 1.000 | 0.714 | 2e-5 | 2e-16 |
| acc90/semi70 | 4.956 | 115.0 | 570 | 0.099 (100/1012) | 1.000 | 0.476 | 1e-6 | 2e-18 |
| acc85/semi90 | 4.920 | 112.0 | 551 | 0.093 (94/1012) | 1.000 | 0.244 | 4e-8 | 1e-20 |
| acc90/semi90 | 4.361 | 123.9 | 540 | 0.071 (72/1012) | 0.999 | 9e-4 | 1e-14 | 2e-30 |
| acc95/semi70 | 4.296 | 136.0 | 584 | 0.083 (84/1012) | 1.000 | 0.037 | 9e-11 | 8e-25 |
| -value at budget | ||||||||
|---|---|---|---|---|---|---|---|---|
| configuration | TPF | forwards | tokens | ( ) | ||||
| acc85/semi70 | 4.253 | 17.2 | 73 | 0.178 (118/664) | 1.000 | 1.000 | 0.978 | 0.081 |
| acc85/semi90 | 3.810 | 18.0 | 69 | 0.145 (96/664) | 1.000 | 1.000 | 0.372 | 1e-4 |
| acc90/semi70 | 3.492 | 19.4 | 68 | 0.139 (92/664) | 1.000 | 0.999 | 0.222 | 2e-5 |
| acc90/semi90 | 3.349 | 20.6 | 69 | 0.087 (58/664) | 1.000 | 0.153 | 9e-7 | 1e-15 |
| acc95/semi70 | 2.980 | 23.7 | 71 | 0.081 (54/664) | 1.000 | 0.059 | 7e-8 | 3e-17 |
| -value at budget | ||||||||
|---|---|---|---|---|---|---|---|---|
| configuration | TPF | forwards | tokens | ( ) | ||||
| acc85/semi70 | 5.973 | 131.7 | 787 | 0.057 (29/512) | 0.789 | 3e-4 | 3e-11 | 2e-20 |
| acc85/semi90 | 5.850 | 129.8 | 759 | 0.070 (36/512) | 0.983 | 0.012 | 2e-8 | 2e-16 |
| acc90/semi70 | 5.317 | 146.2 | 778 | 0.064 (33/512) | 0.941 | 0.003 | 2e-9 | 5e-18 |
| acc90/semi90 | 5.245 | 150.9 | 792 | 0.053 (27/512) | 0.659 | 8e-5 | 3e-12 | 1e-21 |
| acc95/semi70 | 4.534 | 165.4 | 750 | 0.051 (26/512) | 0.584 | 4e-5 | 1e-12 | 2e-22 |
| -value at budget | ||||||||
|---|---|---|---|---|---|---|---|---|
| configuration | TPF | forwards | tokens | ( ) | ||||
| acc85/semi70 | 7.619 | 28.6 | 218 | 0.143 (60/420) | 1.000 | 0.998 | 0.372 | 0.001 |
| acc85/semi90 | 7.224 | 31.8 | 230 | 0.121 (51/420) | 1.000 | 0.936 | 0.055 | 1e-5 |
| acc90/semi70 | 6.536 | 32.7 | 214 | 0.117 (49/420) | 1.000 | 0.887 | 0.030 | 4e-6 |
| acc90/semi90 | 6.339 | 34.2 | 217 | 0.090 (38/420) | 1.000 | 0.290 | 2e-4 | 7e-10 |
| acc95/semi70 | 5.492 | 40.1 | 220 | 0.093 (39/420) | 1.000 | 0.349 | 3e-4 | 2e-9 |
| Redline | MEAN ( ) just slower | MEAN ( ) at or faster | MEAN ( ) | ||||||||
| gain | exc. | ref. | gain | exc. | gain | exc. | gain | exc. | |||
| LLaDA2 math | |||||||||||
| 0.05 | 0.0 | 0.1 | 99.9 | none slower | 0 | 0.6 | 5.7 | 37.5 | 99.6 | ||
| 0.10 | 36.6 | 0.1 | 0.0 | 5 | 36.6 | 0.1 | 6 | 37.3 | 0.1 | 37.5 | 0.1 |
| 0.15 | 37.5 | 0.0 | 0.0 | 6 | 37.3 | 0.0 | 7 | 37.5 | 0.0 | 37.5 | 0.0 |
| 0.20 | 37.5 | 0.0 | 0.0 | 6 | 37.3 | 0.0 | 7 | 37.5 | 0.0 | 37.5 | 0.0 |
| NETHOLM | MEAN ( ) | ||||||||||||
| measure | Redline | BONF | FIXSEQ | UNCORR | PLUGIN | H | HB | WSR | MEANCI | slower | faster | ||
| [0pt][0pt] LLaDA2 math ( ) | |||||||||||||
| 0.05 | gain | 0.0 | 0.0 | 0.1 | 0.2 | 7.8 | 0.0 | 0.0 | 8.6 | 27.0 | none | 0.6 | 37.5 |
| test | 0.1 | 0.1 | 1.6 | 2.4 | 58.4 | 0.0 | 0.0 | 47.9 | 89.6 | 5.7 | 99.6 | ||
| pool | 0.1 | 0.1 | 1.6 | 2.4 | 58.4 | 0.0 | 0.0 | 50.3 | 99.5 | 5.7 | 100.0 | ||
| 0.10 | gain | 36.6 | 32.7 | 36.7 | 37.0 | 37.5 | 0.0 | 0.0 | 37.5 | 37.5 | 36.6 | 37.3 | 37.5 |
| in sample | splits | held out | |||||
| deployed | gain | net | ref. | same | gain | net | |
| LLaDA2 math | |||||||
| 0.05 | reference | +0.0 | 0.0 | 999 | 999 | +0.0 | 0.0 |
| 0.10 | acc85/semi70 | +37.5 | 4.0 | 0 | 894 | +37.6 [+34.1, +41.3] | 4.2 [ 5.9, 2.6] |
| 0.15 | acc85/semi70 | +37.5 | 4.0 | 0 | 1000 | +37.6 [+34.1, +41.4] | 4.0 [ 5.9, 2.2] |
| 0.20 | acc85/semi70 | +37.5 | 4.0 | 0 | 1000 | +37.6 [+34.1, +41.4] | 4.0 [ 5.9, 2.2] |
| deployed | valid | TPF | TPF gain (%) | net accuracy (pp) | |||
|---|---|---|---|---|---|---|---|
| [0pt][0pt] The six configurations shared with the main grid, | |||||||
| 0.05 | reference | 1 of 6 | 0 | 0.000 | 4.42 | ||
| 0.10 | acc85/semi70/add0.10 | 6 of 6 | 53 | 0.052 | 5.90 | ||
| 0.15 | acc85/semi70/add0.10 | 6 of 6 | 53 | 0.052 | 5.90 | ||
| 0.20 | acc85/semi70/add0.10 | 6 of 6 | 53 | 0.052 | 5.90 | ||
| [0pt][0pt] The add slice, | |||||||
| line of work | guarantee | multiplicity | mechanism | generality |
|---|---|---|---|---|
| BD-LM families and serving engine (BD3-LM, SDAR, LLaDA2.0 and 2.1, MBD-LM) | ✗mean accuracy | ✗hand-picked points | stall documented, unexplained | ✗ |
| learned-frontier line (distillation, multi-block visibility, RL shaping, in-place revision, Fast-dLLM and v2, D2F, AdaBlock, dParallel, LoPA, d3LLM, DMax) | ✗ | ✗ | ✗ | one cost each |
| Fast-dLLM Theorem 1 | deterministic, conditional on trusting model confidence, no | ✗ | ✗ | one threshold |
| CALM (LTT on AR early exit) | ✓finite-sample, relative to the full model | one threshold | ✗ | early exit only |
| LTT and conformal risk control (frameworks) | ✓ | ✓procedure | not a serving system | |
| lossy AR acceleration (speculative decoding, Medusa typical acceptance, PTQ) | ✗supply the settings | ✗ | ✗ | one cost each |