Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a natural remedy, yet existing MoE serving systems target autoregressive decoding and rely on intra-iteration layer-wise prefetching: while computing one layer, they predict and load experts for subsequent layers. Under dLLM inference, block-wise routing expands the active expert working set within each iteration, making such prefetches difficult to complete in time and costly when mispredicted. Consequently, existing prefetch-based solutions often degenerate into on-demand expert loading with high decoding latency. We propose OLED-MoE, an expert offloading system that shifts the optimization target from intra-iteration prefetching to inter-iteration expert retention. Its key insight is that adjacent denoising iterations exhibit strong expert routing overlap, and token confidence indicates which experts are likely to be reused. OLED-MoE uses confidence-guided inter-iteration prediction to retain high-value experts in GPU memory without introducing extra prefetch traffic. It further compensates unavoidable cache misses through CPU-GPU cooperative execution, jointly considering dynamic expert computation load and predicted future reuse. Across diverse dLLM workloads, OLED-MoE reduces time per output token (TPOT) by 1.23x-7.93x and improves expert cache utilization by 1.44x-4.23x over state-of-the-art offloading systems. Notably, OLED-MoE approaches full-residency performance while using only 40% of the expert GPU memory, incurring merely 23% higher TPOT despite a 60% reduction in expert memory footprint. OLED-MoE's source code is publicly available at https://github.com/flashserve/OLED-MoE.
Figure 4 . Expert activation footprint and expert FLOPs of semi-AR dLLM decoding compared with AR decoding. Results are measured on LLaDA2.0-mini with block length 16.
Figure 5 . TPOT and expert cache utilization of prior offloading systems under the same GPU expert-cache budget.
Figure 6 . Failure patterns of prior expert offloading systems under MoE-based dLLM workloads: (a) exposed prefetch latency, (b) exposed CPU compensation latency, (c) locality-unaware eviction, and (d) reuse-agnostic placement.
Figure 7 . Inter-iteration locality in dLLMs: the left shows hidden-state similarity across iteration gaps, the right shows expert routing overlap across iteration gaps and the middle shows expert routing overlap in block-boundary transitions.
Figure 8 . System Overview of OLED-MoE.
Figure 9 . Expert priority mapping from token confidence and gate scores to cache-retention priorities.
Figure 10 . Design signals used by OLED-MoE: (a) token confidence and routing stability, (b) confidence-stability distribution, and (c) layer-wise activation ratios.
Figure 11 . Dynamic load- and reuse-aware dispatch: high-reuse experts are transferred first, then bounded swaps balance CPU/GPU latency.
Figure 12 . Average TPOT and expert cache utilization of OLED-MoE and baseline systems.
Figure 13 . Gap between OLED-MoE and the ideal cache bound across expert-cache capacities.
Figure 14 . Prefill latency breakdown of OLED-MoE and FineMoE with cache=100.
Figure 15 . Ablation analysis of OLED-MoE: (a) incremental TPOT reduction from the three design components; (b) detailed comparisons of components.
Figure 16 . Effects of decoding configurations on expert offloading: (a) TPOT, expert activation ratio, and expert cache utilization of OLED-MoE across batch sizes; (b) TPOT of OLED-MoE across block lengths and expert-cache capacities, with circles marking the minimum TPOT under each capacity.
Figure 17 . Sensitivity of OLED-MoE to (a) GPU expert-cache budget and (b) PCIe bandwidth.
Figure 18 . Generality and runtime overhead of OLED-MoE: (a) TPOT on LLaDA2.0 and LLaDA2.1 under cache sizes 40 and 100; (b) policy overhead per forward.
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance while naturally supporting parallel decoding. However, as dLLMs are increasingly integrated with Mixture-of-Experts (MoE) architectures to scale model capacity, a fundamental mismatch arises between block parallel decoding and token-level expert selection. Specifically, each dLLM forward pass processes multiple tokens with bidirectional dependencies, whereas conventional MoE layers route each token independently. This mismatch substantially increases the number of uniquely activated experts, making inference increasingly memory-bound. To address this, we propose dMoE, a simple yet effective block-level MoE framework. The central idea of dMoE is to aggregate token-level expert distributions within each block into a unified block-level expert distribution, which is then used to guide expert routing in a more coherent manner. In this way, dMoE substantially reduces the number of uniquely activated experts during inference without sacrificing performance, thereby mitigating the memory-bound bottleneck. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of dMoE. On average, dMoE reduces the number of uniquely activated experts from 69.5 to 14.6 while retaining 99.11% of the original performance. Meanwhile, it reduces memory usage by 76.64% to 79.84% and achieves 1.14× to 1.66× end-to-end latency speedup. Code is available at: https://github.com/fscdc/dMoE
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization and bidirectional context through parallel block-level decoding. However, as dLLMs continue to scale up with mixture-of-experts (MoE) architectures, their deployment on resource-constrained devices remains an open challenge. Existing AR-based methods often incur either prohibitive I/O overhead or significant compute bottlenecks. In this work, we propose TIDE, a novel resource-efficient inference system that leverages the temporal stability of expert activations during the diffusion process within the block. Specifically, we leverage the temporal stability of expert activations during the diffusion process within the block and introduce an interval-based expert refresh strategy that updates the expert placement in an I/O-aware fashion. To ensure optimal performance, we formulate the inference scheduling as a mathematical programming problem, solving for the optimal interval that minimizes I/O traffic and CPU computation. Most importantly, TIDE is a lossless optimization that requires no model training, providing a "free lunch" acceleration for dLLM inference. In a single GPU-CPU system, we demonstrate that TIDE achieves up to 1.4× and 1.5× throughput improvements over prior baselines on LLaDA2.0-mini and LLaDA2.0-flash models, respectively.
The Mixture of Experts (MoE) architecture has become a fundamental building block in state-of-the-art large language models (LLMs), improving domain-specific expertise in LLMs and scaling model capacity without proportionally increasing their computational overhead. However, MoE inference often suffers from suboptimal GPU utilization, load imbalance, and elevated latency arising from multiple tokens waiting on the same experts for their computation which arises from sparsity of expert activation. To address these challenges, we propose a dynamic expert replication strategy that predicts which experts are likely to be overloaded and replicates them for upcoming batches of tokens. The replicated experts process batch tokens concurrently across layers, which leads to improved parallelism, shorter GPU idle time, and significantly faster inference. Experimental evaluations conducted on large-scale MoE models, including Switch-base-128 and Switch-base-256, demonstrate that our method achieves near-complete GPU utilization (approx 100%), leading to upto 3x improvement in inference speed while preserving approximately 90-95% of the performance of baseline architectures
Ankit Jyothish, Ali Jannesari, Aishwarya Sarkar +1