Organizations: Wuhan University · Shanghai Jiao Tong University · The Hong Kong University of Science and Technology · Damen Database Co., Ltd. · Central China Normal University
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to 2.50× over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to 1.92× for online chatbots.
Figure 2: Block-autoregressive generation with KV refresh.
Figure 3: Per-request latency shares of inference stages (bars: mean ± std; markers: min/max).
Figure 4: Linear fits for prefill and KV-refresh costs.
Figure 5: Architecture of the DWS predictor. An encoder extracts the representations of prompt x and aggregates into z(x) , followed by two separate heads for the output-block and denoising-step distributions. Then the corresponding survival probabilities are combined to reconstruct the DWS D . DWS shown on the right is a smoothed visualization of D for the first prompt in LMSYS-Chat-1M.
Figure 6: Cost-model fidelity across request cost ranges. DWS-denoising isolates the denoising component from our full model, enabling a direct comparison with scalar baselines. Lower is better.
6-8 Predictor
Cost MAE (ms) ↓
Total-step MAE ↓
Per-block step MAE ↓
Block-count MAE ↓
Throughput on CPU (req/s) ↑
Output Length
887.97
–
–
2.01
37.67
Total Steps
517.04
62.50
–
–
37.54
Direct Cost ( untransferable )
432.29
–
–
–
37.55
DWS
355.24
56.06
4.41
2.14
36.35
DWS ( w/o Stage A)
604.43
87.91
6.02
3.68
36.34
DWS ( single -horizon)
445.65
59.70
5.33
2.82
36.51
Table 1: Prediction accuracy of prompt-only predictors and DWS ablations . Cost MAE measures total-cost error; the others measure denoising-structure error. Throughput measurements use the same 256 prompts. “–” denotes a non-applicable metric. Best and second-best results are marked.
Figure 7: Online chatbot serving performance. We report average E2E latency and TTFT (P90 in Appendix I ). All predictors are trained on LMSYS-Chat-1M, so (b) tests cross-dataset generalization .
\rectanglecolor gray!154-25-9 \rectanglecolor cyan!66-27-9 \rectanglecolor blue!58-29-9 Serving Model
Dataset
SDG Metrics
FCFS
PARS (GPU)
PARS-Cost (GPU)
PARS-Cost (CPU)
DWS-SJF (CPU) (Ours)
LLaDA 2.0-mini
(i)
LMSYS- Chat-1M
Time to 1K Req. (s) ↓
328
275
168
290
120
Completed in 3 min ↑
517
661
1021
625
1299
(ii)
Alpaca
Time to 1K Req. (s) ↓
197
159
108
165
75
Completed in 3 min ↑
912
1127
1545
1069
1836
LLaDA 2.0-flash
(iii)
LMSYS- Chat-1M
Time to 1K Req. (s) ↓
1598
1194
1085
1382
592
Completed in 3 min ↑
103
155
181
138
310
Table 2: Performance of schedulers on SDG tasks. All predictors are trained on LMSYS-Chat-1M traces collected with LLaDA2.0-mini. We evaluate (i) in-domain, (ii) cross-dataset , (iii) cross-model , and (iv) cross-model-and-dataset settings to test the generalization of these schedulers.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Request and denoising execution
x
Input prompt of a request.
B
Configured block size, i.e., the maximum number of tokens and denoising steps within one output block.
Nmax
Maximum number of output blocks allowed by the serving configuration.
m
Number of prompt blocks.
n
Number of generated output blocks.
Appendix
Table 3: Summary of notation.
Figure 8: Latency breakdown under varying block counts.
Figure 9: Linear fits for prefill and KV-refresh costs.
Metric
First step
Last step
Δ
Forward latency (ms)
6.043
8.622
+2.579
Routed-expert layers, total (ms)
3.038
5.670
+2.632
Attention layers, total (ms)
1.114
1.116
+0.002
Activated experts per layer
32.7
66.3
+33.6
Padded rows per layer
611
1,102
+491
Appendix
Table 5: Paired comparison between the first and last denoising steps within the same block on LLaDA2.0-mini.
Model
≤ 1K
1–2K
2–4K
4–8K
8–16K
16–32K
αm
8.99%
5.01%
4.18%
3.03%
2.00%
1.61%
α1m+α2m2
7.42%
3.73%
2.50%
2.34%
1.47%
1.52%
Appendix
Table 6: Prefill-cost modeling under long prompts. MAPE of the linear model used in the main paper and a quadratic alternative across different prompt-length ranges.
Model
≤ 1K
1–2K
2–4K
4–8K
8–16K
16–32K
βn
5.71%
6.05%
5.53%
6.60%
8.08%
10.17%
β1n+β2n2
4.15%
4.55%
4.82%
5.70%
7.68%
9.40%
Appendix
Table 7: KV-refresh-cost modeling under long outputs. MAPE of the linear model used in the main paper and a quadratic alternative across different output-length ranges.
Model
≤ 1K
1–2K
2–4K
4–8K
8–16K
16–32K
Wb,s+δ(m−mˉ)
3.29%
2.32%
1.83%
2.03%
2.51%
3.57%
Wb,s+δ1(m−mˉ)+δ2(m−mˉ)2
3.57%
2.95%
1.53%
1.49%
1.87%
2.08%
Appendix
Table 8: Denoising-cost modeling under long prompts. MAPE of the linear prompt-length correction and a quadratic alternative across different prompt-length ranges.
Output length
≤ 1K
1–2K
2–4K
4–8K
8–16K
16–32K
MAPE
3.56%
3.14%
3.40%
4.86%
5.29%
5.73%
Appendix
Table 9: Denoising-cost modeling under long outputs. MAPE of the profiled cell-cost model across different output-length ranges.
Precision
Total-step MAE ↓
Per-block step MAE ↓
Block-count MAE ↓
Throughput (req/s) ↑
FP32 ( Original )
56.06
4.41
2.14
36.35
FP16
56.19
4.47
2.16
17.08
INT8
56.42
4.62
2.19
84.29
INT4
56.92
4.82
2.26
58.71
Appendix
Table 10: Effect of predictor precision. All variants use the same trained DWS predictor and are evaluated on a single CPU core. Lower is better for MAE, while higher is better for throughput.
Figure 10: Characteristics of LLM serving workloads. Top: prompt-length distributions of LMSYS-Chat-1M, ShareGPT, and Alpaca. Bottom: individual service-cost distributions of 4K requests from each dataset served by LLaDA2.0-mini. Dashed and dotted lines denote the mean and median, respectively.
Dataset
Scheduler
Statistic
E2E Latency (s)
TTFT (s)
Per-token Latency (s/token)
LMSYS- Chat-1M
FCFS
Avg
287.93
258.57
5.52
P90
674.90
646.80
15.01
MiniLM-based PARS-Cost (INT8)
Avg
237.12
205.28
2.73
P90
608.73
568.39
3.54
DWS (INT8)-SJF (Ours)
Avg
149.71
116.44
0.58
P90
605.08
556.41
1.07
Appendix
Table 11: Additional online serving metrics at RPS=7 (the maximum RPS evaluated in the main-text online experiments). MiniLM-based PARS-Cost (INT8) uses the same encoder and CPU resources as our predictor. We report average and P90 E2E latency, TTFT, and per-token latency. Serving model is LLaDA2.0-mini. Lower is better.
Serving Model
Dataset
Time to 1K Req. (s) ↓
Completed in 3 min ↑
FCFS
DWS-SJF (CPU) (Ours)
Oracle -SJF
FCFS
DWS-SJF (CPU) (Ours)
Oracle -SJF
LLaDA2.0-mini
LMSYS-Chat-1M
328
120
107
517
1299
1373
Alpaca
197
75
62
912
1836
1959
LLaDA2.0-flash
LMSYS-Chat-1M
1598
592
531
103
310
350
Alpaca
909
407
365
198
468
509
Appendix
Table 12: Comparison with Oracle-SJF on SDG tasks. Oracle-SJF schedules requests by their actual inference costs measured before, serving as an approximate upper bound for SJF scheduling. All predictors are trained on LMSYS-Chat-1M traces collected with LLaDA2.0-mini.
Encoder
Backbone Params. (M) ↓
Cost MAE (ms) ↓
Throughput on CPU (req/s) ↑
MiniLM-L6
23 ( 0.70× )
452.60 ( 1.27× )
70.72 ( 1.95× )
MiniLM-L12 ( Ours )
33 ( 1.00× )
355.24 ( 1.00× )
36.35 ( 1.00× )
BERT-base-uncased
110 ( 3.33× )
338.09 ( 0.95× )
10.85 ( 0.30× )
Qwen3-0.6B ( Yang et al., 2025 )
600 ( 18.18× )
324.71 ( 0.91× )
2.20 ( 0.06× )
Appendix
Table 13: Effect of predictor encoder on DWS prediction. All variants use the same predictor architecture and training setup, differing only in the encoder, and are evaluated without quantization on a single CPU core. Values in parentheses are relative to MiniLM-L12 ( Ours ); green and red denote favorable and unfavorable changes, respectively.
Configuration
Source deployment
Target deployment
GPU
RTX PRO 6000
RTX A6000
GPU architecture
Blackwell
Ampere
GPU memory
96 GiB
48 GiB
GPU power limit
600 W
300 W
Tensor parallelism
TP=1
TP=2
PCIe
PCIe 5.0 × 16
PCIe 4.0 × 16
Appendix
Table 14: Source and target deployment configurations. The two deployments differ substantially in both hardware and system stack. The same trained DWS predictor is reused without any retraining or fine-tuning ; only the deployment-specific cost factors are re-profiled on the target deployment.
Workload
Metric
FCFS
DWS (INT8)-SJF
Improvement
Online / LMSYS-Chat-1M
Avg. E2E latency (s) ↓
1350.14
827.82
1.63×
Avg. TTFT (s) ↓
1288.99
767.36
1.68×
Online / ShareGPT
Avg. E2E latency (s) ↓
1679.76
1105.02
1.52×
Avg. TTFT (s) ↓
1606.41
1033.89
1.55×
Offline / LMSYS-Chat-1M
Time to 1K Req. (s) ↓
678.26
250.73
2.71×
Completed in 3 min ↑
229
747
3.26×
Appendix
Table 15: Serving performance after cross-deployment transfer. The DWS predictor is transferred unchanged from the source deployment, while only the deployment-specific cost profile is refreshed. Online results are reported at RPS=8, the heaviest load evaluated on the target deployment. Offline experiments submit 3K requests following the main-text protocol.
Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead to significant throughput advantages and superior GPU utilization over the traditional autoregressive paradigm. However, this parallelism is constrained by the requirement of a fixed-size response length prior to generation. This architectural limitation imposes a severe trade-off: oversized response length results in computational waste on semantically meaningless padding tokens, while undersized response length causes output truncation requiring costly re-computations that introduce unpredictable latency spikes. To tackle this issue, we propose Predict-then-Diffuse, a simple and model-agnostic framework that enables compute-budgeted inference per input query by first estimating the response length and then using it to run inference with D-LLM. At its core lies an Adaptive Response Length Predictor (AdaRLP), which estimates the optimal response length given an input query. As a measure against under-estimating the response length and re-running inference with a higher value, we introduce a data-driven safety mechanism based on a small increase of the predicted length. As a whole, our framework avoids wasting computation on padding tokens, at the same time preserving output quality. Experimental validation on multiple datasets demonstrates that Predict-then-Diffuse significantly reduces computational costs (FLOP) compared to the default D-LLM inference mechanism, while being robust to skewed data distributions.
Michael Rottoli, Subhankar Roy, Stefano Paraboschi
Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attention prevents exact autoregressive-style KV caching, since committing one position shifts the KV activations of all others. Approximate caching techniques such as Fast-dLLM and dKV-Cache refresh KV activations repeatedly and reuse them across intervening decodes, inducing a repeated prefill/decode structure. This makes AR serving mechanisms relevant to dLLMs, but not directly applicable. dLLM decodes are block-sized rather than token-sized, prefills recur, and bidirectional attention precludes the chunked prefill mechanism used for stall-free colocated serving. We present Sangam, a serving system for cached dLLM inference. Sangam introduces a deficit token-budget scheduler that admits in-flight decodes first, admits whole indivisible prefills only when the accumulated token budget allows, and carries unused budget forward. This achieves amortized stall-free scheduling. Disaggregated serving avoids prefill-decode interference but suffers from prefill/decode resource partitioning problem. Sangam adopts a hybrid serving strategy, overflowing prefills onto decode workers to relieve prefill under-provisioning, and uses the same deficit-budget scheduler to protect those workers' decodes from the overflow. We show that like AR serving, dLLM serving design space is governed by prefill-decode interference and prefill/decode partitioning. Colocated serving is most effective on decode-heavy workloads, cutting mean latency by 9-20% over hybrid execution on LLaDA-8B ShareGPT; while hybrid execution is most effective on prefill-heavy workloads, cutting mean latency by 8-20% over colocated execution on Dream-7B arXiv. Sangam is available at https://github.com/UT-InfraAI/sangam.
Nitin Kedia, Saurabh Agarwal, Myungjin Lee +1
The University of Texas at Austin Austin, United States · Cisco Research Bellevue, United States
Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency. We present BlockServe, a continuous batching framework that integrates block-grained scheduling -- immediately evicting completed requests at block boundaries -- with mixed-state execution that extends dual cache and parallel decoding to heterogeneous batches via gather-scatter indexing. Furthermore, a compute-aware admission controller expands effective batch capacity through token-budgeted refill. On Dream and LLaDA across five benchmarks, BlockServe achieves 1.9--10.6× throughput over Fast-dLLM with comparable generation quality, establishing block-grained scheduling as a foundation for high-throughput offline dLLM inference.