Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.
Figures & tables
System
Compat.
Evolv.
Deriv.
Prog.
NCCL [ 36 ]
✓
✗
✗
✓
NVSHMEM [ 20 ]
✓
✗
✗
✓
MSCCL++ [ 14 ]
✓
✓
✗
✓
NCCLX [ 57 ]
✓
✗
✗
✗
ParallelKittens [ 48 ]
✗
✗
✗
✓
Purlin
✓
✓
✓
✓
Table 1 : Properties of scale-up communication systems. Compatibility : supports Ampere, Hopper, and Blackwell, as well as older GPUs that rely on vectorized loads and stores. Evolvability : extends the datapath while retaining collective semantics and orchestration. Derivability : derives orchestration from collective semantics without an explicit communication schedule. Programmability : exposes interfaces for customizing primitives and collective implementations. ✓ denotes hardware extensions accompanied by algorithm changes (Evolv.), or programmable primitives with limited collective customization (Prog.).
Figure 1 : Purlin’s decoupled stack. Collective layouts and codesign policies describe what to do; the SNAC layer orchestrates; the Atom deals only with how bytes move on a given GPU generation. Codesign policies tune execution to improve performance.
Parallelism
Representative collectives
Role
Tensor [ 47 ]
AllReduce
Combine partial results
Expert [ 59 ]
AllGather(V), ReduceScatter(V), AllToAll(V)
Dispatch tokens and combine expert outputs
Data-parallel attention [ 59 ]
AllGather(V), ReduceScatter(V)
Exchange activations between attention and FFN or MoE layers
Ulysses sequence [ 15 ]
AllToAll
Redistribute sequence and attention-head partitions
Table 2 : Collective communication in distributed inference. (V) denotes a variable-length variant.
Figure 2 : Latency breakdown for Qwen3.5-122B-A10B on eight A100 GPUs with TP8/EP8/DP4 and 1K input/output tokens. Rows show p50 end-to-end latency, p50 Time Per Output Token (TPOT), and p50 Time To First Token (TTFT). Communication accounts for 24–32% of SGLang’s latency across these metrics.
GPU
Bandwidth ( R )
Latency ( τ )
Rτ
Rounded BDP
A100
300 GB/s
2.2 µs
644 KiB
512 KiB
H200
450 GB/s
2.1 µs
936 KiB
1 MiB
B200
900 GB/s
2.8 µs
2.4 MiB
2 MiB
Table 3 : BDP estimates using nominal unidirectional bandwidth and measured round-trip NVLink latency. The last column is rounded to the closest power of 2. Latency results obtained from Figure 7 .
Figure 3 : Bandwidth for a 256 MiB pull over NVLink with 256 threads per SM. We tune NCCL’s device API. Dotted lines mark nominal peak bandwidth per direction. Purlin reaches comparable or higher bandwidth with fewer SMs.
Figure 4 : Layouts on three ranks. Top: ReduceScatter combines partition k from every input into yk at rank k ; AllGather places all yk in rank order at every rank. Bottom: AllToAll gives rank 1 partition 1 from every source. Colors identify the input partition index. These layouts describe the required results independently of any transfer mechanism.
Figure 6 : Execution timeline of composed allReduce . All ranks stage chunks incrementally. Here, rank k reduces chunks from partition k from all ranks while also gathering reduced chunks of other partitions from other ranks
Figure 7 : Push and pull over NVLink through Atom::copy , using 256 threads per CTA. Dotted lines mark nominal peak bandwidth per direction. Each transfer involves two GPUs.
Component
Lines of code
Generic Atom
187
Ampere specialization
299
Hopper specialization
296
Blackwell specialization
53
Shared types and utilities
611
Table 4: Atom implementations, including local helpers.
Figure 9 : Collective performance on ordinary application buffers on eight B200, H200 and A100 GPUs.
GPU and interconnect
LLM checkpoint (precision)
Total parameters
A100-SXM4 (80 GB) NVLink 3
Qwen3.5-122B-A10B [ 42 ] (BF16)
122B
H200 (141 GB) NVLink 4
DeepSeek-V4-Flash [ 43 ] (FP8)
291B
B200 (180 GB) NVLink 5
DeepSeek-V4-Pro [ 35 ] (NVFP4)
1.6T
Table 5: Evaluation platforms and LLM configurations.
Figure 10 : AllGather and AllReduce on eight B200 GPUs with peer-accessible buffers. Purlin-ZS omits staging. Lower is better for latency and higher is better for bandwidth. In the table, we count launched CTAs and bound concurrent SM use by min(CTAs launched,148) at 1 GiB. NCCL Symm AllGather uses the copy engine and launches no kernel.
Figure 11 : Copy bandwidth for one-sided transfers between two GPUs as we increase the Atom’s pipeline depth from one to eight stages, for both pull and push. Messages are 256 MiB and each CTA has 256 threads. We hold the CTA count fixed at eight on A100 and H200 and sixteen on B200. Dotted lines mark nominal per-direction link bandwidth.
Figure 12 : Application performance on eight GPUs. (a) Offline LLM serving sweeps concurrency through 1, 4, 16, 64, and 256. Rows use input/output lengths of 1,000/1,000 tokens (chat), 8,000/1,000 (summary), and 1,000/8,000 (reasoning). Interactivity is the inverse of median TPOT and throughput is output tokens per second per GPU. Higher and farther right is better. (b) Online serving replays the Mooncake conversation trace at target rates from 0.2 to 3.4 requests/s. The upper row shows the full sweep on logarithmic axes and the lower row expands low-load behavior. (c) Qwen-Image uses Ulysses sequence parallelism and 50 denoising steps. The lower row expands resolutions from 1282 to 10242 . Panels (b) and (c) report median latencies, where lower is better.
Metric
SGLang
+ Purlin
AIME26: 30 problems, 16 samples each
pass@1 (%, mean ± SEM)
96.04±0.34
95.62±0.40
pass@16 (%)
96.67
96.67
majority@16 (%)
96.67
96.67
Qwen-Image: 100 prompts, matched seeds
Mean ImageReward
1.2063
1.2063
Table 6: Application output quality on eight H200 GPUs. We compare DeepSeek-V4-Flash FP8 on AIME26 and Qwen-Image at 10242 resolution. LLM sampling is unseeded, whereas image seeds match between systems. SEM denotes standard error of the mean.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 14 : AllGather and AllReduce on eight A100 GPUs with peer-accessible buffers. The left column reports latency and the right column reports algorithm bandwidth. Purlin-ZS omits staging.
Figure 15 : AllGather and AllReduce on eight H200 GPUs with peer-accessible buffers, using the same layout and timing method as Figure 14 . Lower is better for latency (left column) and higher is better for bandwidth (right column).
Figure 16 : Variable-length collective performance with ordinary application buffers on eight B200, H200, and A100 GPUs. All systems use a Zipf partition with exponent s=0.125 . Message size is the total T defined in Eq. 4 .
Collective / baseline
A100
H200
B200
ReduceScatter / NCCL Symm
1.48 / 1.25
1.19 / 0.93
1.61 / 1.24
AllToAll / ParallelKittens
—
1.26 / 0.92
2.14 / 1.02
Appendix
Table 7 : Additional peer-accessible results on eight GPUs. Each cell reports geometric-mean latency speedup / bandwidth ratio for Purlin-ZS relative to the named baseline; values above one favor Purlin-ZS. The sweeps cover 1 KiB–512 KiB and 1 MiB–1 GiB, respectively. ParallelKittens does not support A100.
Batch inference has become a central mode of AI computation, yet existing inference engines still rely on execution models designed for interactive serving. When scaled to millions of sequences, batch workloads reveal two fundamental requirements: the ability to handle extreme inter- and intra-sequence load variation that emerges only at runtime, and the ability to sustain high utilization across large fleets of GPUs. Existing systems fail to meet these requirements, losing substantial fractions of achievable throughput. We introduce a new architectural foundation for batch inference: the sequence coroutine compute model, which represents each sequence as a fine-grained, event-driven coroutine. This model exposes expressive primitives that allow the runtime to reorganize work dynamically, enabling larger expert-level batches, mitigating stragglers, reallocating work across devices, and maintaining utilization even on cost-effective or memory-constrained GPUs. Building on this abstraction, we implement BatchGen, a production-ready system that uses the coroutine model at cluster scale. On a 128-GPU cluster, BatchGen reduces batch completion time by up to 2.3×, and on memory-constrained accelerators it outperforms the strongest offloading baseline by up to 9.6×. We will open-source BatchGen at https://github.com/batchgen-project/batchgen
Training Large Language Models (LLMs) on heterogeneous clusters presents significant challenges for collective communication, as hardware from multiple vendors introduces diverse network and computational characteristics. Existing collective communication frameworks (e.g., NCCL, RCCL) designed for homogeneous environments fail to address mixed-hardware setups, while communication libraries with heterogeneous support (e.g., Gloo, OpenMPI) incur heavy overhead in the data path. This paper presents HetCCL, a framework that enables heterogeneous collective communication by efficient P2P transport across heterogeneous devices (e.g., GPUs), eliminating the host-device memory copy overhead while offloading the control to the CPUs. For combining collectives (e.g., AllReduce, ReduceScatter), HetCCL introduces a border-communicator mechanism that achieves vendor independence by using the intrinsic reduction in the combining collectives in vendor collective communication libraries. With efficient heterogeneous P2P transport and portable reduction mechanism, HetCCL proposes a hierarchical topology abstraction for heterogeneous clusters, dissecting collective communication into cluster-level primitives that guarantee optimal cross-cluster data transfer volume and optimal bandwidth utilization. We implement HetCCL with 4 different vendor support and evaluate it in 4 heterogeneous settings with benchmarks and end-to-end LLM tasks. Our evaluation shows that HetCCL achieves 17-19x higher bandwidth than Gloo in heterogeneous communications, and speeds up end-to-end training by up to 16.9% in the per-step-time.
Yuejie Wang, Tao Chang, Yuanyuan Zhao +10
Peking University · Beijing Academy of Artificial Intelligence · Infrawaves +1
The rapid growth of large language models (LLMs) has made GPU communication a critical bottleneck. While prior work reduces communication volume via quantization or lossy compression, these approaches introduce numerical errors that can degrade convergence, accuracy, and stability. We present UCCL-Zip, a unified design that integrates lossless compression directly into GPU communication primitives. UCCL-Zip supports both point-to-point (P2P) and collective communication without modifying user-facing APIs or compromising numerical correctness. For P2P communication, Uzip-P2P employs a split-send pipeline that exposes transmissible data early and overlaps compression with communication, while preserving high GPU efficiency by operating on large data blocks. For collective communication, Uzip-NCCL integrates compression into NCCL's persistent kernel model via fused execution, eliminating redundant memory traffic and kernel launches. In real workloads, UCCL-Zip accelerates RL weight synchronization by up to 47.5% and reduces vLLM end-to-end inference latency by up to 10%, all without application changes.
Shuang Ma, Chon Lam Lao, Zhiying Xu +8
UC Davis · Harvard University · Amazon Web Services +1