Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a natural remedy, yet existing MoE serving systems target autoregressive decoding and rely on intra-iteration layer-wise prefetching: while computing one layer, they predict and load experts for subsequent layers. Under dLLM inference, block-wise routing expands the active expert working set within each iteration, making such prefetches difficult to complete in time and costly when mispredicted. Consequently, existing prefetch-based solutions often degenerate into on-demand expert loading with high decoding latency. We propose OLED-MoE, an expert offloading system that shifts the optimization target from intra-iteration prefetching to inter-iteration expert retention. Its key insight is that adjacent denoising iterations exhibit strong expert routing overlap, and token confidence indicates which experts are likely to be reused. OLED-MoE uses confidence-guided inter-iteration prediction to retain high-value experts in GPU memory without introducing extra prefetch traffic. It further compensates unavoidable cache misses through CPU-GPU cooperative execution, jointly considering dynamic expert computation load and predicted future reuse. Across diverse dLLM workloads, OLED-MoE reduces time per output token (TPOT) by 1.23x-7.93x and improves expert cache utilization by 1.44x-4.23x over state-of-the-art offloading systems. Notably, OLED-MoE approaches full-residency performance while using only 40% of the expert GPU memory, incurring merely 23% higher TPOT despite a 60% reduction in expert memory footprint. OLED-MoE's source code is publicly available at https://github.com/flashserve/OLED-MoE.
Figure 4 . Expert activation footprint and expert FLOPs of semi-AR dLLM decoding compared with AR decoding. Results are measured on LLaDA2.0-mini with block length 16.
Figure 5 . TPOT and expert cache utilization of prior offloading systems under the same GPU expert-cache budget.
Figure 6 . Failure patterns of prior expert offloading systems under MoE-based dLLM workloads: (a) exposed prefetch latency, (b) exposed CPU compensation latency, (c) locality-unaware eviction, and (d) reuse-agnostic placement.
Figure 7 . Inter-iteration locality in dLLMs: the left shows hidden-state similarity across iteration gaps, the right shows expert routing overlap across iteration gaps and the middle shows expert routing overlap in block-boundary transitions.
Figure 8 . System Overview of OLED-MoE.
Figure 9 . Expert priority mapping from token confidence and gate scores to cache-retention priorities.
Figure 10 . Design signals used by OLED-MoE: (a) token confidence and routing stability, (b) confidence-stability distribution, and (c) layer-wise activation ratios.
Figure 11 . Dynamic load- and reuse-aware dispatch: high-reuse experts are transferred first, then bounded swaps balance CPU/GPU latency.
Figure 12 . Average TPOT and expert cache utilization of OLED-MoE and baseline systems.
Figure 13 . Gap between OLED-MoE and the ideal cache bound across expert-cache capacities.
Figure 14 . Prefill latency breakdown of OLED-MoE and FineMoE with cache=100.
Figure 15 . Ablation analysis of OLED-MoE: (a) incremental TPOT reduction from the three design components; (b) detailed comparisons of components.
Figure 16 . Effects of decoding configurations on expert offloading: (a) TPOT, expert activation ratio, and expert cache utilization of OLED-MoE across batch sizes; (b) TPOT of OLED-MoE across block lengths and expert-cache capacities, with circles marking the minimum TPOT under each capacity.
Figure 17 . Sensitivity of OLED-MoE to (a) GPU expert-cache budget and (b) PCIe bandwidth.
Figure 18 . Generality and runtime overhead of OLED-MoE: (a) TPOT on LLaDA2.0 and LLaDA2.1 under cache sizes 40 and 100; (b) policy overhead per forward.