Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.
Figures & tables
Figure 1 : Mooncake Architecture.
Figure 2 : Normalized throughput and latency of prefill and decoding stages with different sequence lengths or batch sizes for the dummy LLaMA2-70B model.
Figure 3 : The KVCache pool in CPU memory. Each block is attached with a hash value determined by both its own hash and its prefix for deduplication.
Figure 4 : Workflow of inference instances. ( ∗ ) For prefill instances, the load and store operations of the KVCache layer are performed layer-by-layer and in parallel with the prefill computation to mitigate transmission overhead (see § 5.2 ). ( † ) For decoding instances, asynchronous loading is performed concurrently with GPU decoding to prevent GPU idle time.
Figure 5 : Input and output length distributions in the request trace.
Figure 6 : CDF (Cumulative Distribution Function) of the block hit count in the request trace.
Block capacity
Inf
100000
50000
30000
10000
1000
LRUCache
0.55
0.55
0.55
0.54
0.46
0.34
LFUCache
0.55
0.55
0.55
0.50
0.40
0.34
LengthAwareCache
0.55
0.54
0.53
0.50
0.40
0.34
Table 1: Cache hit rates under different cache policies and capacities.
Figure 7 : Latency of storing KVCache of different request lengths (Layer-wise latency refers to the difference in latency between Layer-wise Prefill and Prefill without storing KVCache).
Figure 8 : The prefill scheduling experiment in the Mooncake cluster.
Figure 9 : The load of prefill and decoding instances over 20 minutes, before using the prediction-based early rejection.
Figure 10 : Instance load when applying Early Rejection and Early Rejection Based on Prediction.
Dataset
Avg Input Length
Avg Output Length
Cache Ratio
Arrival Pattern
ArXiv Summarization [ 26 ]
8088
229
~0%
Poisson Process
L-Eval [ 27 ]
19019
72
>80%
Poisson Process
Simulated Data
16k, 32k, 64k, 128k
512
50%
Poisson Process
Real Data
7955
194
~50%
Timestamp-based
Table 2: Datasets used in the end-to-end experiment.
Figure 11 : End-to-end experiments of Mooncake and vLLM on the ArXiv Summarization and L-Eval datasets
Figure 12 : End-to-end experiments of Mooncake and vLLM on simulated data.
Figure 13 : Request TTFT and TBT distributions of Mooncake and vLLM under real workloads
Baseline
Early Rejection
Early Rejection based on Prediction
Number of rejected requests
4183
3771
3589
Table 3: Number of requests rejected by the system under the overloaded-scenario experiment.
School of Computing and Information Systems, The University of Melbourne · School of Computer Science and Technology, Huazhong University of Science and Technology (www.ruizhang.info)