Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.
Figures & tables
System
Compat.
Evolv.
Deriv.
Prog.
NCCL [ 36 ]
✓
✗
✗
✓
NVSHMEM [ 20 ]
✓
✗
✗
✓
MSCCL++ [ 14 ]
✓
✓
✗
✓
NCCLX [ 57 ]
✓
✗
✗
✗
ParallelKittens [ 48 ]
✗
✗
✗
✓
Purlin
✓
✓
✓
✓
Table 1 : Properties of scale-up communication systems. Compatibility : supports Ampere, Hopper, and Blackwell, as well as older GPUs that rely on vectorized loads and stores. Evolvability : extends the datapath while retaining collective semantics and orchestration. Derivability : derives orchestration from collective semantics without an explicit communication schedule. Programmability : exposes interfaces for customizing primitives and collective implementations. ✓ denotes hardware extensions accompanied by algorithm changes (Evolv.), or programmable primitives with limited collective customization (Prog.).
Figure 1 : Purlin’s decoupled stack. Collective layouts and codesign policies describe what to do; the SNAC layer orchestrates; the Atom deals only with how bytes move on a given GPU generation. Codesign policies tune execution to improve performance.
Parallelism
Representative collectives
Role
Tensor [ 47 ]
AllReduce
Combine partial results
Expert [ 59 ]
AllGather(V), ReduceScatter(V), AllToAll(V)
Dispatch tokens and combine expert outputs
Data-parallel attention [ 59 ]
AllGather(V), ReduceScatter(V)
Exchange activations between attention and FFN or MoE layers
Ulysses sequence [ 15 ]
AllToAll
Redistribute sequence and attention-head partitions
Table 2 : Collective communication in distributed inference. (V) denotes a variable-length variant.
Figure 2 : Latency breakdown for Qwen3.5-122B-A10B on eight A100 GPUs with TP8/EP8/DP4 and 1K input/output tokens. Rows show p50 end-to-end latency, p50 Time Per Output Token (TPOT), and p50 Time To First Token (TTFT). Communication accounts for 24–32% of SGLang’s latency across these metrics.
GPU
Bandwidth ( R )
Latency ( τ )
Rτ
Rounded BDP
A100
300 GB/s
2.2 µs
644 KiB
512 KiB
H200
450 GB/s
2.1 µs
936 KiB
1 MiB
B200
900 GB/s
2.8 µs
2.4 MiB
2 MiB
Table 3 : BDP estimates using nominal unidirectional bandwidth and measured round-trip NVLink latency. The last column is rounded to the closest power of 2. Latency results obtained from Figure 7 .
Figure 3 : Bandwidth for a 256 MiB pull over NVLink with 256 threads per SM. We tune NCCL’s device API. Dotted lines mark nominal peak bandwidth per direction. Purlin reaches comparable or higher bandwidth with fewer SMs.
Figure 4 : Layouts on three ranks. Top: ReduceScatter combines partition k from every input into yk at rank k ; AllGather places all yk in rank order at every rank. Bottom: AllToAll gives rank 1 partition 1 from every source. Colors identify the input partition index. These layouts describe the required results independently of any transfer mechanism.
Figure 6 : Execution timeline of composed allReduce . All ranks stage chunks incrementally. Here, rank k reduces chunks from partition k from all ranks while also gathering reduced chunks of other partitions from other ranks
Figure 7 : Push and pull over NVLink through Atom::copy , using 256 threads per CTA. Dotted lines mark nominal peak bandwidth per direction. Each transfer involves two GPUs.
Component
Lines of code
Generic Atom
187
Ampere specialization
299
Hopper specialization
296
Blackwell specialization
53
Shared types and utilities
611
Table 4: Atom implementations, including local helpers.
Figure 9 : Collective performance on ordinary application buffers on eight B200, H200 and A100 GPUs.
GPU and interconnect
LLM checkpoint (precision)
Total parameters
A100-SXM4 (80 GB) NVLink 3
Qwen3.5-122B-A10B [ 42 ] (BF16)
122B
H200 (141 GB) NVLink 4
DeepSeek-V4-Flash [ 43 ] (FP8)
291B
B200 (180 GB) NVLink 5
DeepSeek-V4-Pro [ 35 ] (NVFP4)
1.6T
Table 5: Evaluation platforms and LLM configurations.
Figure 10 : AllGather and AllReduce on eight B200 GPUs with peer-accessible buffers. Purlin-ZS omits staging. Lower is better for latency and higher is better for bandwidth. In the table, we count launched CTAs and bound concurrent SM use by min(CTAs launched,148) at 1 GiB. NCCL Symm AllGather uses the copy engine and launches no kernel.
Figure 11 : Copy bandwidth for one-sided transfers between two GPUs as we increase the Atom’s pipeline depth from one to eight stages, for both pull and push. Messages are 256 MiB and each CTA has 256 threads. We hold the CTA count fixed at eight on A100 and H200 and sixteen on B200. Dotted lines mark nominal per-direction link bandwidth.
Figure 12 : Application performance on eight GPUs. (a) Offline LLM serving sweeps concurrency through 1, 4, 16, 64, and 256. Rows use input/output lengths of 1,000/1,000 tokens (chat), 8,000/1,000 (summary), and 1,000/8,000 (reasoning). Interactivity is the inverse of median TPOT and throughput is output tokens per second per GPU. Higher and farther right is better. (b) Online serving replays the Mooncake conversation trace at target rates from 0.2 to 3.4 requests/s. The upper row shows the full sweep on logarithmic axes and the lower row expands low-load behavior. (c) Qwen-Image uses Ulysses sequence parallelism and 50 denoising steps. The lower row expands resolutions from 1282 to 10242 . Panels (b) and (c) report median latencies, where lower is better.
Metric
SGLang
+ Purlin
AIME26: 30 problems, 16 samples each
pass@1 (%, mean ± SEM)
96.04±0.34
95.62±0.40
pass@16 (%)
96.67
96.67
majority@16 (%)
96.67
96.67
Qwen-Image: 100 prompts, matched seeds
Mean ImageReward
1.2063
1.2063
Table 6: Application output quality on eight H200 GPUs. We compare DeepSeek-V4-Flash FP8 on AIME26 and Qwen-Image at 10242 resolution. LLM sampling is unseeded, whereas image seeds match between systems. SEM denotes standard error of the mean.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 14 : AllGather and AllReduce on eight A100 GPUs with peer-accessible buffers. The left column reports latency and the right column reports algorithm bandwidth. Purlin-ZS omits staging.
Figure 15 : AllGather and AllReduce on eight H200 GPUs with peer-accessible buffers, using the same layout and timing method as Figure 14 . Lower is better for latency (left column) and higher is better for bandwidth (right column).
Figure 16 : Variable-length collective performance with ordinary application buffers on eight B200, H200, and A100 GPUs. All systems use a Zipf partition with exponent s=0.125 . Message size is the total T defined in Eq. 4 .
Collective / baseline
A100
H200
B200
ReduceScatter / NCCL Symm
1.48 / 1.25
1.19 / 0.93
1.61 / 1.24
AllToAll / ParallelKittens
—
1.26 / 0.92
2.14 / 1.02
Appendix
Table 7 : Additional peer-accessible results on eight GPUs. Each cell reports geometric-mean latency speedup / bandwidth ratio for Purlin-ZS relative to the named baseline; values above one favor Purlin-ZS. The sweeps cover 1 KiB–512 KiB and 1 MiB–1 GiB, respectively. ParallelKittens does not support A100.