Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands. Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78× speedup and 2.99--3.93× higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55× speedup and 6.07× higher energy efficiency over Jetson AGX Orin.
Figures & tables
Fig. 1: (a) Iterative decoding process of a dLLM. (b) Fast-dLLM’s part-wise decoding with approximate KV-cache reuse.
Fig. 2: Three inefficiencies in dLLM decoding and DynaTE’s corresponding algorithm–hardware optimizations.
Fig. 3: Workload characterization of the three optimization opportunities: (a) per-token update gain across denoising iterations; (b) oracle local-conditioning gain between adjacent unresolved tokens; (c) full-logit HBM traffic; and (d) bursty per-token versus smoothed multi-token processing demand.
Fig. 4: Token-wise pruning dataflow. Query-skip removes selected tokens from query-side attention and FFN computation while retaining their keys and values for the remaining tokens; no-skip executes the complete layer.
Fig. 5: Overall architecture of DynaTE.
Fig. 6: (a) Reconfigurable PE array, (b) internal PE structure, and (c) multi-precision splittable MAC unit.
Fig. 7: Representative PE-array reconfiguration modes: (a) canonical dataflow and (b)–(c) reshaped dataflows using left-buffer bypass and inter-subarray transfers.
Fig. 8: Streaming vocabulary-processing engine, including output reshaping, burst-to-stream smoothing, multi-token Top-32 selection, and shared distribution-statistics generation.
Fig. 9: Multi-token Top-32 engine and its interleaved tournament-pipeline schedule.
Fig. 10: Merge–Split–Merge dataflow for integrating local serial refinement into parallel dLLM decoding.
Parameter
Configuration
Technology / Frequency
TSMC 28 nm / 1 GHz
PE Total MACs
2,048
Array configurations
64×32 – 512×4
Weight–activation precision
W8A8 or W8A4
Activation buffer
2×128 KB
Weight buffer
2×64 KB, double-buffered
TABLE I: DynaTE hardware configuration.
Fig. 11: Survivor-buffer depth exploration.
Fig. 12: Tournament-core count exploration.
Fig. 13: Token-pruning effectiveness across LLaDA-8B and Dream-7B: relative total Transformer MAC count before and after pruning.
Fig. 14: FLDD effectiveness across LLaDA-8B and Dream-7B.
Fig. 15: Dependency locality and FLDD capture: (a) cumulative distribution of absolute token distances among oracle dependency pairs; (b) fraction of the same oracle pairs captured by FLDD.
Method
GSM8K
H.Eval
MBPP
MATH
MMLU
Baseline
70.3 / 74.2
35.4 / 39.5
40.0 / 44.3
31.4 / 29.1
43.9 / 42.2
+ Pruning
69.9 / 74.1
35.2 / 39.3
39.7 / 44.5
30.8 / 28.8
43.6 / 41.6
+ FLDD
70.4 / 74.2
35.3 / 39.4
40.0 / 44.1
31.9 / 29.2
43.8 / 42.0
+ Quant.
70.5 / 73.9
35.2 / 39.2
39.7 / 44.1
31.3 / 28.9
43.8 / 42.3
+ All
69.8 / 73.9
34.9 / 39.4
39.3 / 43.7
30.8 / 28.9
43.7 / 41.8
TABLE II: Ablation study on generation quality. Each entry reports LLaDA-8B / Dream-7B.
Fig. 16: End-to-end comparison across LLaDA-8B and Dream-7B: (a) speedup and (b) energy-efficiency improvement.
Fig. 17: Average PE utilization of the fixed and reconfigurable arrays across LLaDA-8B and Dream-7B workloads.
Fig. 18: Normalized per-iteration latency breakdown with and without the MSM dataflow.
Metric
Logic
SRAM
HBM
Overall
Area (mm 2 )
2.12
1.79
–
3.91
Power (W)
1.01
1.18
5.17
7.36
TABLE III: Area and power breakdown of DynaTE.
Fig. 20: Intermediate-logit traffic and normalized HBM access energy for HBM-resident logits and streaming summaries.
Fig. 21: Top-32 scheduling and filtering effectiveness: (a) tournament-core utilization under single-token and multi-token scheduling; (b) normalized logits admitted to the tournament tree without and with Prior-ID seeding.
Fig. 22: Incremental ablation of DynaTE.
Metric
DART
dLLM-OPU
SA
DynaTE
Platform
ASIC
U200 FPGA
ASIC
ASIC
Technology
7 nm
16 nm
28 nm
28 nm
Frequency
1 GHz
300 MHz
1 GHz
1 GHz
Area (mm 2 )
3.20 a
–
3.34
3.91
Throughput (tokens/s)
8.97
12.12
5.76
24.89
Area eff. (tokens/s/mm 2 )
2.81
–
1.72
6.36
TABLE IV: Comparison with prior dLLM accelerators.
*Equal contribution. The work is done during Tianyi’s internship at CBG Celia DeviceAI Team, Huawei. · 1Harbin Institute of Technology, Shenzhen · 2Huawei +1