Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.
Figures & tables
Figure 1: Execution timeline and memory usage of pipeline parallel execution (left) and tensor parallel execution (right) using three GPUs.
Figure 2: Overview of the self-speculative model and its execution timeline in Spexis under a two-stage pipeline. The timeline shows the execution of a single request: orange arrows indicate speculative execution, and black arrows indicate normal execution after speculation rejection.
PP
TP
SP
SP α
Tpp(B/N)B/N
Ttp(B)B
Tsp(B)Beff⋆
Tsp(Bα⋆)Beff,α⋆
Table 1: Estimation of throughput for PP, TP, SP and SP α . N denotes the number of pipeline stages, B the maximum batch size, and T∗(⋅) the single-stage decoding time for a given batch size. SP α incorporates selective speculation described in Section 5.1 .
GPU
Interconnect
Model
Thpt. (tokens/s)
Speedup
Spexis
TP
H100 SXM
NVLink via NVSwitch
L70
3587
3652
0.98 ×
Q32
2402
5068
0.47 ×
A100 SXM
NVLink via NVSwitch
L70
1228
966
1.27 ×
Q32
1454
2058
0.71 ×
RTX PRO 6000
PCIe 5.0 × 16
L70
1492
1236
1.21 ×
Table 2: Fixed-batch decoding throughput (tokens/s) of Spexis and TP-only on single-node, eight-GPU Runpod configurations, measured at a fixed batch size of 64. Speedup is computed as Spexis throughput divided by TP throughput; bold indicates the higher throughput.
Model
Exit
Task Accuracy
Avg.
Math
QA
Sum
RAG
Tran
Mult
L70
1/2
.840
.743
.760
.788
.774
.793
.785
1/4
.680
.629
.581
.622
.570
.656
.634
L8
1/2
.800
.776
.660
.727
.644
.759
.759
1/4
.665
.609
.487
.547
.461
.625
.596
Q32
1/2
.682
.535
.565
.619
.609
.603
.596
Table 3: Speculation accuracy on SpecBench for Llama-3.3-70B-Instruct (L70), Llama-3.1-8B-Instruct (L8), Qwen3-32B (Q32), and Qwen3-8B (Q8). The drafting adapter (Exit) is inserted at either the halfway point (1/2) or the one-quarter point (1/4) of the model.
ShareGPT
SpecBench
UltraChat
Model
Thresh.
Prec.
Recall
Prec.
Recall
Prec.
Recall
L70
L<4
79.0%
53.4%
82.6%
59.7%
76.3%
56.0%
L<8
78.6%
55.2%
80.3%
59.5%
78.1%
57.6%
L<16
81.2%
58.1%
85.8%
59.4%
81.5%
61.9%
L<32
83.5%
64.0%
88.9%
61.3%
82.9%
68.6%
L<64
85.2%
67.6%
90.7%
67.5%
85.8%
71.1%
Table 4: Precision (Prec.) and recall for the remaining-length prediction using thresholds of 4, …, 64 tokens (L70: Llama-3.3-70B-Instruct, Q32: Qwen3-32B).
Speculative Decoding (SD) accelerates low-concurrency LLM inference by employing a draft-then-verify paradigm. However, mainstream methods typically rely on multi-token prediction, which introduces escalating prediction difficulty and serial drafting latency. To address these, we propose Speculative Pipeline Decoding (SPD), a groundbreaking framework that unlocks the true potential of pipeline parallelism. By partitioning the target LLM into n pipeline stages, SPD allows LLM to process n tokens within single sequence in parallel to accelerate decoding. To continuous fill the pipeline in single sequence decoding, a speculation module aggregates intermediate features across different pipeline depths to predict the next token, executing strictly in parallel with the target model's pipeline step, to realize bounded difficulty, higher acceptance rates, and zero latency bubbles. Our experiments demonstrate that SPD achieves significantly higher theoretical and wall-clock speedup compared to mainstream baselines at moderate pipeline depth, though more aggressive settings require further improvement. Our code is available at https://github.com/yuyijiong/speculative_pipeline_decoding
Speculative decoding can substantially accelerate LLM inference, but realizing its benefits in practice is challenging due to evolving workloads. We present TIDE (Temporal Incremental Draft Engine), a serving-engine-native framework that integrates online draft adaptation directly into high-performance LLM inference systems. TIDE reuses target model's intermediate hidden states generated during inference as training signals for draft adaptation, thereby avoiding additional target model computation and serving-time overhead. It employs adaptive runtime control to activate speculation and draft model training only when beneficial. TIDE exploits heterogeneous clusters by mapping inference and training to appropriate GPU classes. Across diverse real-world workloads, TIDE achieves up to 1.66× throughput over no-speculation baselines while recovering performance on misaligned workloads where static draft models degrade throughput. TIDE also reduces training time by up to 3.02× and storage requirements by 24× compared to existing draft training approaches, and improves system throughput by up to 1.22× on heterogeneous GPU clusters.
Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a single draft-verify schedule, even though the optimal draft length varies widely across requests and changes dynamically within each request. We propose ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components. First, a unified mixed forward allows drafting and verifying requests to coexist in the same batched forward pass, removing the need for global draft-verify phases. Second, a lightweight online speculation scheduler uses per-request acceptance-rate estimates and a batch-aware cost model to let each request independently choose when to verify. Third, an intra-draft refresh layer performs full attention at a single designated layer during drafting, updating the sparse context at every draft step to reduce staleness during drafting. Across three models and five reasoning and long-context benchmarks, ASPIRE achieves 1.70-4.58× speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately 27% over the strongest prior self-speculative baselines.
Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao +2
University of Southern California, Los Angeles, CA, USA