Organizations: Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518052, China · University of Chinese Academy of Sciences, Beijing 100049, China · Victoria University of Wellington, Wellington, New Zealand · Institute of AI and Brain Sciences, Department of Computer Science, University of Macau, Macau 999078, China
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash
Figures & tables
Fig. 1: Three connected challenges in DFlash serving. (1) A shared prefill and verify tile leaves 16 live and 112 masked rows per request and head. (2) Equal 16-slot windows accommodate unequal accepted prefixes. (3) Changed boundaries make old slices read R1 and miss R2. Equal-width padding restores 16 slots per request, while CPU re-slicing delays verification. The dashed arrow links Challenges 2 and 3. Token lengths are illustrative.
Fig. 2: DFlash progress on 306,215 request-steps from a mixed-dataset training-split trace collection at c=8 . Progress includes one anchor. The dashed line marks the mean, 6.09. Prefixes to the right of the width-8 boundary lose their excess tokens under uniform truncation.
Fig. 3: Fixed verification widths on Qwen3-8B and GSM8K, A100, TP1, temperature 0, c=8 , with full-length DFlash drafting and no predictor. Each width uses one run with 8 warmup and 256 measured requests. Left shows output-token throughput. Right averages periodically logged progress lengths, including the anchor. Width 16 is native DFlash.
Fig. 4: Draft logit gaps and acceptance on 31,814 Qwen3-8B and 31,188 Qwen3-4B held-out request-steps. Each point gives mean gap and accepted draft tokens in one of ten equal-count bins. Acceptance excludes the anchor. Spearman correlations use all request-steps.
Fig. 5: DScale request flow and module organization. Dashed boxes mark captured graphs, excluding input staging and state commit. The dashed prefill branch shares tile specialization. Lower and right paths return unfinished requests to the scheduler and completed results to users, respectively.
Fig. 6: DVL allocator. The predictor supplies initial lengths and prefix scores to the Budget allocator. The Flat-buffer packer stores prefixes and offsets. Short strips select from full draft blocks. Grow and shrink are alternative budget conditions. Colors identify requests.
Fig. 7: Token throughput at concurrency 8–32. Top and bottom rows show Qwen3-8B and Qwen3-4B. Columns show GSM8K, MATH500, HumanEval, and MBPP.
Fig. 8: End-to-end request latency at concurrency 8–32. Top and bottom rows show Qwen3-8B and Qwen3-4B. Columns show mean, p95, and p99. Each point is the median across four datasets of their five-repetition medians.
Fig. 9: GSM8K cumulative ablation using five-repetition medians, successively adding tile routing, DVL, and Graph integration.
Fig. 10: Accepted-token retention before and after budget adjustment on 31,814 validation request-steps for Qwen3-8B and 31,188 for Qwen3-4B. Features and full-width labels are fixed. Each collection-step cohort of N retained rows receives 8N slots after adjustment.
Fig. 11: Decode-step time, allocator costs, and graph-pool memory under TP1. Panels a–c use GSM8K at temperature 0. Panels a and b show mean step times and reductions from DFlash. Panel c reports c=8 allocator time in milliseconds and as a percentage of DScale step time. Overlapping component intervals count once. The Budget allocator and Flat-buffer packer are timed together. Panel d reports reserved memory in MiB, showing means and sample standard deviations over three independent process starts.
Method
Drafter reused a
Per-request budget b
Compact variable-length verify
CUDA Graph integration c
Extra predictor d
No hardware performance-curve profiling for budget scheduling e
DFlash [ 10 , 2 ]
✓
×
×
✓
×
✓
DSpark [ 11 ]
×
✓
✓
✓
✓
×
Domino [ 20 ]
×
×
×
✓
×
✓
DScale
✓
✓
✓
✓
✓
✓
TABLE I: Verification-side serving capabilities of the compared execution paths.
Block diffusion speculative decoding accelerates LLM inference by predicting all tokens within a block simultaneously for the target model to verify in parallel. Predicting an entire block at once requires a sufficiently capable draft model and effective utilization of the target model's internal knowledge. However, the state-of-the-art method DFlash constrains all draft layers to share a single fused representation derived from only a few target layers, limiting per-layer expressiveness and hindering further scaling of draft capacity. In this paper, we present \modelname, which flares out the narrow conditioning bottleneck of DFlash through a lightweight layer-wise fusion mechanism: each draft layer attends to its own learnable combination of a broad set of target layers at negligible overhead, simultaneously injecting richer target knowledge and providing every draft layer with a distinct input. This enhanced per-layer expressiveness enables scaling the draft model to deeper architectures with consistent gains. We further scale training data from 800K to 2.4M samples to fully exploit the enlarged capacity. On six benchmarks spanning mathematical reasoning, code generation, and conversation, \modelname attains average wall-clock speedups of 5.52x on Qwen3-4B, 5.46x on Qwen3-8B, and 3.91x on GPT-OSS-20B, improving over DFlash by roughly 11%, 8%, and 5% respectively. Our code is available at https://github.com/Tencent/AngelSlim.
Jiebin Zhang, Zhenghan Yu, Song Liu +9
School of Computer Science, Peking University · Tencent
Diffusion language models (dLLMs) generate text by iteratively denoising multiple token positions in parallel, offering an attractive alternative to strictly autoregressive decoding. In practice, however, block-wise dLLM inference exposes a difficult granularity trade-off: small blocks preserve local conditioning but require many denoising steps, whereas large blocks expose more parallelism but can make premature commitments and accumulate cache error. Existing acceleration methods typically choose a single block size per request, leaving the complementarity among block sizes unused. We show that block size itself is a useful branching dimension. Different block sizes induce related but non-identical KV-cache trajectories: branches often share an initial prefix, bifurcate at semantically decisive positions, and later agree on syntactically lightweight tokens. Motivated by this structure, we propose BlockBatch, a training-free online inference framework that executes multiple block-size branches for the same request inside a batched forward pass. BlockBatch coordinates these branches through confidence-gated token merging, leader-based synchronization, and periodic full-sequence refreshes that re-anchor local block updates to a globally consistent KV state. Across 3 representative dLLMs and 4 datasets, BlockBatch reduces denoising NFEs by 26.6% on average and achieves a 1.33× average end-to-end speedup over Fast-dLLM while preserving accuracy. These results identify block-size diversity as a practical and previously underexplored axis for branch-parallel dLLM inference.
Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass. Recent diffusion-based drafters generate an entire block of tokens in parallel but usually commit to a single draft sequence per verification: once the first mismatch occurs, all subsequent draft tokens are discarded, resulting in a limited acceptance rate. Naively batching more draft candidate sequences only introduces a marginal improvement, as redundant or poorly placed branches increase the cost of drafting and verification without proportionally increasing the number of accepted tokens. We propose D^2SD, a dual diffusion draft speculative decoding framework that organizes candidates into a confidence-guided prefix tree, where the first diffusion drafter generates a block along with per-position confidence scores that are used to identify the most likely rejection boundary and select the top-K prefix ranges for recovery; the second variable-prefix diffusion drafter re-anchors at each selected prefix and proposes alternative continuations in one batched pass; the resulting shared-prefix candidates are jointly verified via cascade attention. Empirically, D^2SD shows clear improvements over both the underlying diffusion approach and strong autoregressive speculative decoding baselines.
Liyuan Zhang, Jiarui Zhang, Jinwei Yao +6
Peking University · Tsinghua University · HKUST +2