Recently, Mixture-of-Experts (MoE) models have gained attention for efficiently scaling large language models. Although these models are extremely large, their sparse activation enables inference to be performed by accessing only a fraction of the model at a time. This property opens the possibility of on-device inference of MoE, which was previously considered infeasible for such large models. Consequently, various systems have been proposed to leverage this sparsity and enable efficient MoE inference for edge devices. However, previous MoE inference systems like Fiddler[8] or DAOP[13] rely on DRAM-based offloading and are not suitable for memory constrained on-device environments. As recent MoE models grow to hundreds of gigabytes, RAM-offloading solutions become impractical. To address this, we propose FlashMoE, a system that offloads inactive experts to SSD, enabling efficient MoE inference under limited RAM. FlashMoE incorporates a lightweight ML-based caching strategy that adaptively combines recency and frequency signals to maximize expert reuse, significantly reducing storage I/O. In addition, we built a user-grade desktop platform to demonstrate the practicality of FlashMoE. On this real hardware setup, FlashMoE improves cache hit rate by up to 51% over well-known offloading policies such as LRU and LFU, and achieves up to 2.6x speedup compared to existing MoE inference systems.
Figures & tables
Figure 1 . Illustration of an MoE model decoder layer.
Figure 2 . Heatmap comparing LRU and Belady cache policies. The deepest color means the experts are routed at that time step, and the color changes lighter as time goes backward. Evictions based on Belady policy is marked as red ’X’ shape, while LRU is marked as green ’O’. Due to LRU’s recency-based replacement policy, it fails to evict experts immediately after routing as the optimal algorithm does. Additionally, frequently accessed experts are sometimes evicted just before routing, highlighting a key limitation of LRU. This indicates the need for incorporating frequency-based metrics such as LFU and leveraging strategies.
Figure 3 . (a) Overall system of FlashMoE. (b) FlashMoE’s prefill process.
Figure 4 . Overall training method illustration of FlashMoE.
Figure 5 . Decoder layer computation pipeline of FlashMoE.
Hyperparameter
Value
Loss
MSELoss
Optimizer
AdamW
Learning Rate
1e-3
Learning Rate Decay
1e-2
# Layer
3
Hidden size
128
Table 1. Training hyperparameters for FlashMoE
Figure 6 . Hit rate by cache size & replacement policy (a): OLMoE-1B-7B, (b): Qwen3-30B-A3B.
Table 2. Specifications of Desktop System used for FlashMoE.
Figure 7 . (a,d) Model load & prefill latency of the MoE inference system. (b,e) Token generation speed under different cache replacement policies. (c,f) Token generation speed across different inference systems. ((a,b,c): OLMoE-1B-7B, (d,e,f): Qwen3-30B-A3B)