DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory
Authors: Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung, Thanh Tuan Dao
Organizations: Department of Next Generation Semiconductor Convergence and Open Sharing System (COSS), Seoul National University, Seoul 08826, South Korea · Van-Lang Institute of Semiconductor Technology (VIST), Hanoi, Vietnam · Efficient Computation Research Group, Moreh Vietnam · Moreh
Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this design by allowing a single stored model to support both full-accuracy and lower-precision execution, making the effective weight footprint runtime-dependent. This creates an opportunity under bursty workloads, where temporary spikes in KV-cache demand often determine throughput and SLO compliance. We present DPS, a dual-precision LLM serving system that turns weight memory into an elastic resource: under normal load, DPS serves the full-accuracy model; under KV pressure, it switches to a nested, lower-precision variant and repurposes unused weight memory for KV cache blocks. DPS is built on Semi-Unified Memory (SUM), which partitions the weight region into a persistent lower-precision sub-region and a shared region that alternates between residual weight tensors and KV-cache blocks, preserving compatibility with paged KV-cache management. We implement DPS on top of vLLM and evaluate it across both dense and MoE models and various production workload traces. Our results show that \sysname improves sustained throughput by 2.1--3.3× and effective pass@1 by up to +41,pp over Static FP16, while preserving FP16-class accuracy.
Figures & tables
Model
Method
BBH
GPQA
MATH-500
MMLU Pro
MUSR
IFEval
LiveCodeBench
Avg Δ
Phi-3.5-MoE
BF16
74.00
36.38
37.00
60.09
45.90
64.88
22.84
—
FP8
76.27
35.94
38.80
59.66
45.77
42.70
21.80
−2.88
DPS full mode
73.91
35.94
36.20
60.31
46.03
66.73
23.70
+0.25
DPS fast mode
74.11
35.49
37.60
59.73
47.35
65.62
22.94
+0.25
Qwen3-30B-A3B
BF16
56.40
43.53
75.80
72.42
42.46
81.70
56.21
—
FP8
57.32
44.87
73.40
72.54
41.27
82.07
56.02
−0.15
TABLE I: Offline benchmark accuracy among BF16, FP8, and DPS with FP16-only and FP8-only. The rightmost column reports the mean per-benchmark difference relative to BF16 (in percentage points). Per-benchmark variation of ± 1–2 pp is within the noise floor of these benchmarks.
Time in mode (%)
Transitions
Restore
Model
Full
Perf
F → P
P → F
Latency (ms)
Phi-3.5-MoE
100.0
0.0
0
0
–
Qwen3-30B-A3B
20.8
79.2
2
2
24651
GLM-4.7-Flash
52.9
47.1
2
2
8497
TABLE II: Mode switching overhead under BurstGPT window’s traces. Restore latency is the per-cycle mean of the weight H2D wall time.
Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy. However, in KV-cache-intensive serving scenarios, MoEs often exhibit a tension between the GPU memory requirements of the model weights and the growing KV cache. We propose PagedWeight, a novel management method for MoE LLM serving that dynamically quantizes MoE model's weights at runtime and balances expert-weight precision with the KV cache sizes. PagedWeight exposes and effectively navigates the complex tradeoff between the model's task accuracy, memory consumption, and throughput/latency. Across several memory-sensitive MoE serving scenarios, PagedWeight improves the quality-memory tradeoff over several existing quantization baselines. PagedWeight achieves FP16-equivalent accuracy with up to 72.0% GPU memory savings and 1.94× throughput improvement, and improves quality over quantization methods by up to 39.3% at a similar memory budget with at most 4.1% throughput loss.
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly prefill the same content. As the retrieval budget grows, time-to-first-token (TTFT) increases even though the underlying memory is reused across requests. We present InferScale, a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state. InferScale precomputes each memory fact's KV representation, stores it alongside a semantic embedding on the GPU, retrieves relevant facts at serving time, and injects their KV directly into vLLM's paged cache. To support dynamically assembled memories under rotary position embeddings, we introduce Chunked RoPE, which stores keys before rotation and applies their serving-time positions during injection. However, encoding memory facts independently omits the cross-fact context available during joint prefilling. We mitigate this with Context-Window Encoding, which encodes each memory fact together with a small window of preceding conversation context while caching only the target fact's KV. InferScale is implemented through vLLM's KV-connector interface, requiring neither engine modifications nor model fine-tuning. Across three open-weight models on LoCoMo, InferScale keeps TTFT nearly constant as the retrieval budget increases: at k=50 it reduces TTFT by 72-79% (3.6-4.8x), achieves 60.3% accuracy versus 63.3% for Mem0 without serving-time recomputation, and delivers 3.7-4.5x the throughput under concurrent load. Reusable KV state thus decouples memory-conditioned serving latency from retrieved-context size while preserving application quality.
Peter Li, Prashant Pandey
Northeastern University Boston, Massachusetts, USA
Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable and model-determined, while KV-cache is transient and demand-determined. Because cold models rarely reach peak KV-cache demand at the same time, reserving worst-case KV capacity per model wastes memory; a shared KV-cache pool can instead provision aggregate active demand. However, KV-cache sharing is not sufficient when weights and KV-cache remain in a monolithic GPU memory pool. Static weights compete with dynamic KV-cache, and KV-head-limited attention under cold, low-concurrency traffic exposes only a fraction of replicated KV capacity, leading to low GPU memory utilization and weak long-context support. We present CrossPool, a serving engine for cold MoE models that separates FFN weights and KV-cache into two GPU memory pools: a weights pool that consolidates FFN weights across cold models, and a KV-cache pool that dynamically serves active requests while keeping attention local to KV-cache. CrossPool combines a KV-cache planner and virtualizer, a layer-wise pipeline scheduler that hides hidden-state transfers, and persistent kernels with control lowering to reduce CPU-GPU control overhead. With efficient GPU memory pooling, CrossPool underpins bursty long-context requests and outperforms the state-of-the-art kvcached-based multi-LLM serving system, reducing P99 TBT by up to 10.4x.