Organizations: State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Layout
Weights
KV stored
KV read
Attention work
TP
WA/T
B
B
B/T
CP
WA
B/T
B/T
B/T
DP-attention
WA
⌈B/T⌉
⌈B/T⌉
⌈B/T⌉
DOP
WA/T
⌈B/T⌉
⌈B/T⌉
⌈B/T⌉
Appendix
Table 2: Busiest-rank load for B equal-length requests on T ranks, with count-balanced request ownership and ideal token partitioning for CP. KV stored and KV read count histories of ks bytes, and attention work counts one request’s attention. Equation 10 converts these loads into time. The owner entries equal B/T whenever T divides B , matching Equation 1 .
Device / fraction
Layout
GiB/card
GiB/group
Distinct tokens
B200 / 0.85
TP
59.84
478.72
1,045,312
DOP
59.19
473.51
8,271,360
CP
48.22
385.79
6,738,944
DP-attention
46.50
371.98
6,497,792
H200 / 0.90
TP
36.21
289.68
632,512
DOP
33.86
270.90
4,731,904
Appendix
Table 3: KV pools and effective capacity for accelerator groups. The platforms use different models: GLM-5.3 on B200, DeepSeek-V3.2 on H200, and GLM-5.3-Flash on BW1000 DCUs. Memory fraction is the configured allocation fraction. Pool sizes are in GiB; effective capacity counts distinct tokens across the group, excluding replicated copies. Displayed GiB values are rounded.
Input
B
TP
CP
DP-attn.
DOP
Input-length sweep, B=16
1,024
16
1.529
1.176
1.000
1.053
4,096
16
3.285
3.188
3.110
3.111
16,384
16
7.362
8.539
9.380
8.738
32,768
16
9.279
13.723
13.910
12.594
65,536
16
11.140
17.806
18.424
16.287
Appendix
Table 4: B200 fixed-layout sweeps, in total throughput ( 103 tokens/s). Each row has B requests at client concurrency B and 1,024 output tokens per request. Input lengths are exact token counts. DP-attn. abbreviates DP-attention.
Input
B
TP
CP
DP-attn.
DOP
Input-length sweep, B=16
1,024
16
1.234
1.139
0.985
1.110
4,096
16
2.706
2.632
2.339
2.593
16,384
16
5.301
7.005
7.337
7.122
32,768
16
6.999
10.559
11.515
10.989
65,536
16
8.366
14.283
16.205
15.248
Appendix
Table 5: Updated DCU fixed-layout sweeps, in total throughput ( 103 tokens/s). Each row has B requests at client concurrency B and 1,024 output tokens per request. Input lengths are exact token counts. DP-attn. abbreviates DP-attention.
Input
B
TP
CP
DP-attn.
DOP
Input-length sweep, B=16
1,024
16
0.937
0.885
0.930
0.864
4,096
16
1.994
2.080
2.232
2.071
16,384
16
5.037
6.267
6.441
5.996
32,768
16
6.300
10.512
10.137
9.185
65,536
16
7.700
13.279
11.820
11.676
Appendix
Table 6: H200/DeepSeek-V3.2 fixed-layout sweeps, in total throughput ( 103 tokens/s). Each row has B requests and 1,024 output tokens per request. Input lengths are exact token counts. DP-attn. abbreviates DP-attention.
LLM serving optimization typically benchmarks many configurations and reaches for heavy profilers when latency targets are missed. We argue for the reverse discipline: estimation is the analytical layer of profiling -- without it, optimization degenerates to grid search. Floor First is a residual-driven triage workflow. Each decode step is modeled as a five-dimensional resource vector (HBM bytes, FLOPs, network bytes, network messages, KV capacity); summing within a resource and maximizing across resources gives an optimistic floor, the plain sum a pessimistic one. Where a measurement lands inside this [max, sum] interval reads out overlap quality before any profiler is opened, and profilers escalate only on residuals above a stated threshold. Deployment alternatives are compared by wall ordering -- which resource wall binds first as load grows -- rather than by point benchmarks. The account is compositional: new attention or state-space variants enter by declaring one module, and the workflow ships as a zero-dependency calculator plus an agent skill that enforces the discipline in agentic optimization loops. As a case study we analyze a DeepSeek-V3.2-style 671B MoE/MLA model on 16 NVIDIA H20 GPUs, whose ridge point of ~74 FLOP/byte (vs ~590 for H100) makes it an extreme decode-oriented part. The floors show TP16 decoding is KV-capacity-limited to ~70 concurrent 8K requests; sparse attention removes the KV-bandwidth term but not the capacity wall; an EP16+DP-attention layout accepts slightly worse same-batch weight traffic for an order-of-magnitude higher capacity wall (~644) -- while single-stream latency favors TP by 2.4x. The layout judgment is thus a computable function of the operating point, explaining why production deployments on identical hardware have shipped opposite attention layouts.
Long-context LLM serving is bottlenecked by the cost of attending over ever-growing KV caches. Dynamic sparse attention promises relief by accessing only a small, query-dependent subset of the KV state per decoding step and extending the KV storage to CPU memory. In practice, however, these algorithmic savings rarely translate into end-to-end system-level gains because sparse methods typically operate at different granularities and thus rely on ad hoc, per-algorithm implementations. At the same time, hierarchical KV storage introduces a new systems bottleneck: retrieving fine-grained, irregular KV subsets across the GPU-CPU boundary can easily erase the benefits of sparsity. We present SPIN, a sparse-attention-aware inference framework that co-designs the execution pipeline with hierarchical KV storage through three techniques: (1) a unified partition abstraction that maps different sparsity granularities onto a shared page-based KV substrate; (2) a locality-aware KV cache manager that dynamically sizes per-request HBM budgets and uses a GPU-friendly bucketed LRU policy to cut PCIe round-trips; and (3) a two-level hierarchical metadata layout sized to the active working set rather than the worst-case address space. Built on vLLM with three representative sparse attention algorithms, SPIN delivers 1.66-5.66x higher end-to-end throughput and 7-9x lower TTFT than vLLM, and reduces TPOT by up to 58% over the original sparse-attention implementations.
Zihan Zhao, Baotong Lu, Shengjie Lin +8
University of Virginia · Microsoft Research · Georgia Institute of Technology +1
LLM serving exhibits extreme length variability, making size-based scheduling difficult in practice. Recent LLM schedulers approximate SJF/SRPT using predicted decode lengths or ranks and primarily report mean-centric metrics such as TTFT and TBT. We show that these prediction-driven policies can be fragile under distribution shifts, bursty arrivals, and GPU memory pressure, while offering limited control over the tail latency (P90-P99) that dominates user experience, even with perfect decode-length knowledge. We introduce a distribution-aware, prediction-free scheduling framework that replaces explicit length prediction with soft priority boosting driven by lightweight statistical signals. Our design co-optimizes scheduling and cache-aware preemption to account for memory-coupled decode dynamics across workload mixes. Evaluated on production and open-source traces, our method reduces P99 TTLT by up to 35-50% relative to SRPT with perfect length knowledge and reduces TTFT by 34-47% across workloads, including reasoning-heavy and chat-heavy tasks. These results demonstrate a robust alternative for optimizing tail latency in online LLM serving.
Yueying Li, Yuanfan Chen, Jiayang Chen +6
Cornell University, Computer Science Department, NY, USA · Microsoft Azure System Research, WA, USA · Cornell University, Electrical and Computer Engineering Department, NY, USA +2