Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs
Organizations: Office of CTO Dell Technologies Inc. Canada
Abstract
Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selecting the most effective parallelism strategy remains challenging due to complex interactions among computation, communication, pipeline utilization, sequence length, batch size, and model architecture. Existing approaches largely rely on empirical evaluation and provide limited analytical insight into the trade-offs among these strategies, particularly across the distinct prefill and decoding phases of inference. In this paper, we present a unified analytical framework for modeling distributed LLM inference under TP, PP, and HB. The framework decomposes end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead, and derives analytical models that capture TP collective communication, PP point-to-point communication, and pipeline utilization as functions of hardware, model, and workload characteristics. The model further characterizes the differing execution behavior of prefill and decoding, explaining why PP-oriented configurations favor compute-intensive prefill while TP-oriented configurations reduce decoding latency by eliminating pipeline bubbles. Experiments with modern LLMs on multi-GPU platforms validate the model and confirm the fundamental compute-communication trade-off across parallelism strategies. The framework provides practical guidance for parallelism selection, capacity planning, and optimization of future LLM serving systems.
Figures & tables
| Category | Kernel Name |
|---|---|
| GEMM | ampere_bf16_s16816gemm_bf16_256x128_ldg8_relu_f2f_stages_64x3_tn |
| ampere_bf16_s16816gemm_bf16_64x64_ldg8_relu_f2f_stages_64x5_tn | |
| ampere_bf16_s16816gemm_bf16_256x128_ldg8_f2f_stages_64x3_tn | |
| ampere_bf16_s16816gemm_bf16_128x64_ldg8_f2f_stages_64x4_tn | |
| ampere_bf16_s16816gemm_bf16_64x64_sliced1x2_ldg8_f2f_stages_64x6_tn | |
| ampere_bf16_s16816gemm_bf16_64x64_sliced1x2_ldg8_relu_f2f_stages_64x5_tn |
| Parameter | Value |
|---|---|
| GPU Platform | 8 NVIDIA A100 SXM4 (80 GB) |
| GPU Memory | 80 GB HBM2e per GPU |
| HBM Bandwidth | 2.0 TB/s per GPU |
| Peak Compute | 312 TFLOPS (FP16/BF16 Tensor Core) |
| Interconnect | NVLink 600 GB/s |
| Inference Framework | vllm v0.15.1 |