cs.LGSep 30, 2026

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

Authors: Haoyu Zheng, Fangcheng Fu, Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang, Yuanyuan Zhu, Xiao Yan, +1 more

Organizations: Wuhan University · Shanghai Jiao Tong University · The Hong Kong University of Science and Technology · Damen Database Co., Ltd. · Central China Normal University

Abstract

As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to 2.50×2.50\times over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to 1.92×1.92\times for online chatbots.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs

    May 5, 2026Michael Rottoli, Subhankar Roy, Stefano ParaboschiDiffusion Language ModelsLLM Inference Optimization

  2. Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

    Jul 5, 2026Nitin Kedia, Saurabh Agarwal, Myungjin Lee +1Large Language Model ServingDiffusion Language Models

  3. BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

    Jul 9, 2026Yuanjie Zhu, Liangwei Yang, Ke Xu +4Large Language Model ServingSchedulers