Organizations: State Key Laboratory of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Layout
Weights
KV stored
KV read
Attention work
TP
WA/T
B
B
B/T
CP
WA
B/T
B/T
B/T
DP-attention
WA
⌈B/T⌉
⌈B/T⌉
⌈B/T⌉
DOP
WA/T
⌈B/T⌉
⌈B/T⌉
⌈B/T⌉
Appendix
Table 2: Busiest-rank load for B equal-length requests on T ranks, with count-balanced request ownership and ideal token partitioning for CP. KV stored and KV read count histories of ks bytes, and attention work counts one request’s attention. Equation 10 converts these loads into time. The owner entries equal B/T whenever T divides B , matching Equation 1 .
Device / fraction
Layout
GiB/card
GiB/group
Distinct tokens
B200 / 0.85
TP
59.84
478.72
1,045,312
DOP
59.19
473.51
8,271,360
CP
48.22
385.79
6,738,944
DP-attention
46.50
371.98
6,497,792
H200 / 0.90
TP
36.21
289.68
632,512
DOP
33.86
270.90
4,731,904
Appendix
Table 3: KV pools and effective capacity for accelerator groups. The platforms use different models: GLM-5.3 on B200, DeepSeek-V3.2 on H200, and GLM-5.3-Flash on BW1000 DCUs. Memory fraction is the configured allocation fraction. Pool sizes are in GiB; effective capacity counts distinct tokens across the group, excluding replicated copies. Displayed GiB values are rounded.
Input
B
TP
CP
DP-attn.
DOP
Input-length sweep, B=16
1,024
16
1.529
1.176
1.000
1.053
4,096
16
3.285
3.188
3.110
3.111
16,384
16
7.362
8.539
9.380
8.738
32,768
16
9.279
13.723
13.910
12.594
65,536
16
11.140
17.806
18.424
16.287
Appendix
Table 4: B200 fixed-layout sweeps, in total throughput ( 103 tokens/s). Each row has B requests at client concurrency B and 1,024 output tokens per request. Input lengths are exact token counts. DP-attn. abbreviates DP-attention.
Input
B
TP
CP
DP-attn.
DOP
Input-length sweep, B=16
1,024
16
1.234
1.139
0.985
1.110
4,096
16
2.706
2.632
2.339
2.593
16,384
16
5.301
7.005
7.337
7.122
32,768
16
6.999
10.559
11.515
10.989
65,536
16
8.366
14.283
16.205
15.248
Appendix
Table 5: Updated DCU fixed-layout sweeps, in total throughput ( 103 tokens/s). Each row has B requests at client concurrency B and 1,024 output tokens per request. Input lengths are exact token counts. DP-attn. abbreviates DP-attention.
Input
B
TP
CP
DP-attn.
DOP
Input-length sweep, B=16
1,024
16
0.937
0.885
0.930
0.864
4,096
16
1.994
2.080
2.232
2.071
16,384
16
5.037
6.267
6.441
5.996
32,768
16
6.300
10.512
10.137
9.185
65,536
16
7.700
13.279
11.820
11.676
Appendix
Table 6: H200/DeepSeek-V3.2 fixed-layout sweeps, in total throughput ( 103 tokens/s). Each row has B requests and 1,024 output tokens per request. Input lengths are exact token counts. DP-attn. abbreviates DP-attention.
Cornell University, Computer Science Department, NY, USA · Microsoft Azure System Research, WA, USA · Cornell University, Electrical and Computer Engineering Department, NY, USA +2