Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.
Figures & tables
Figure 1: Execution timeline and memory usage of pipeline parallel execution (left) and tensor parallel execution (right) using three GPUs.
Figure 2: Overview of the self-speculative model and its execution timeline in Spexis under a two-stage pipeline. The timeline shows the execution of a single request: orange arrows indicate speculative execution, and black arrows indicate normal execution after speculation rejection.
PP
TP
SP
SP α
Tpp(B/N)B/N
Ttp(B)B
Tsp(B)Beff⋆
Tsp(Bα⋆)Beff,α⋆
Table 1: Estimation of throughput for PP, TP, SP and SP α . N denotes the number of pipeline stages, B the maximum batch size, and T∗(⋅) the single-stage decoding time for a given batch size. SP α incorporates selective speculation described in Section 5.1 .
GPU
Interconnect
Model
Thpt. (tokens/s)
Speedup
Spexis
TP
H100 SXM
NVLink via NVSwitch
L70
3587
3652
0.98 ×
Q32
2402
5068
0.47 ×
A100 SXM
NVLink via NVSwitch
L70
1228
966
1.27 ×
Q32
1454
2058
0.71 ×
RTX PRO 6000
PCIe 5.0 × 16
L70
1492
1236
1.21 ×
Table 2: Fixed-batch decoding throughput (tokens/s) of Spexis and TP-only on single-node, eight-GPU Runpod configurations, measured at a fixed batch size of 64. Speedup is computed as Spexis throughput divided by TP throughput; bold indicates the higher throughput.
Model
Exit
Task Accuracy
Avg.
Math
QA
Sum
RAG
Tran
Mult
L70
1/2
.840
.743
.760
.788
.774
.793
.785
1/4
.680
.629
.581
.622
.570
.656
.634
L8
1/2
.800
.776
.660
.727
.644
.759
.759
1/4
.665
.609
.487
.547
.461
.625
.596
Q32
1/2
.682
.535
.565
.619
.609
.603
.596
Table 3: Speculation accuracy on SpecBench for Llama-3.3-70B-Instruct (L70), Llama-3.1-8B-Instruct (L8), Qwen3-32B (Q32), and Qwen3-8B (Q8). The drafting adapter (Exit) is inserted at either the halfway point (1/2) or the one-quarter point (1/4) of the model.
ShareGPT
SpecBench
UltraChat
Model
Thresh.
Prec.
Recall
Prec.
Recall
Prec.
Recall
L70
L<4
79.0%
53.4%
82.6%
59.7%
76.3%
56.0%
L<8
78.6%
55.2%
80.3%
59.5%
78.1%
57.6%
L<16
81.2%
58.1%
85.8%
59.4%
81.5%
61.9%
L<32
83.5%
64.0%
88.9%
61.3%
82.9%
68.6%
L<64
85.2%
67.6%
90.7%
67.5%
85.8%
71.1%
Table 4: Precision (Prec.) and recall for the remaining-length prediction using thresholds of 4, …, 64 tokens (L70: Llama-3.3-70B-Instruct, Q32: Qwen3-32B).