The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks primarily focus on simple single-turn chatbot workloads. LLM applications are increasingly agentic: coding agents, terminal execution systems, and tool-use agents issue multi-turn requests with growing context lengths. We introduce AgentPerfBench, a benchmark suite for agentic inference. It uses real traces from agentic benchmarks, such as SWE-Bench and TerminalBench, alongside standard chat baselines. This enables benchmarking of models on multi-turn tasks involving tool calling, skill utilization, and increasing context lengths. AgentPerfBench also samples from empirical distributions of input length, output length, and turn count derived from the real traces, generating representative synthetic profiles for cheap and accurate measurements on new hardware. In addition, we further find that several existing benchmarks fail to accurately reflect real hardware performance for two key reasons: 1) they do not account for realistic context-length growth, and 2) they measure inference performance without operating at hardware saturation. We discuss these issues in detail and provide rich kernel-level Nsight Compute (NCU) traces to construct a new multi-dimensional roofline model that captures hardware-system limitations in both memory bandwidth and memory capacity footprint. The benchmarking suite then includes automated scripts to identify potential bottleneck conditions on emerging hardware when evaluated with diverse agentic traces. Together, these contributions quantify the chat-to-agentic gap in current inference benchmarks and characterise per-kernel GPU resource utilisation via roofline analysis.
Figures & tables
Chat
Coding agent
Terminal agent
Computer-use agent
Model
TP
TTFT
TPOT
Req/s
TTFT
TPOT
Req/s
TTFT
TPOT
Req/s
TTFT
TPOT
Req/s
Dense
LLaMA-3.1-8B
1
343
9.4
13.5
11719
227.3
2.4
6889
197.2
3.0
1037
22.8
6.5
Qwen3.5-9B
1
953
12.0
11.0
6954
168.7
3.2
2877
47.1
4.6
1539
22.5
5.9
Qwen3.5-27B
2
1518
21.5
6.0
2729
62.3
5.3
2519
52.0
4.7
2709
33.2
3.6
MoE
Table 4: Medium multi-turn trace_replay saturation metrics on H100-SXM5-80GB with vLLM v0.19.0 and prefix caching enabled. Each cell reports the observed saturation point for one model/profile pair; TTFT and TPOT are medians in ms. Column groups use the workload names from the main text; exact profile IDs are listed in Appendix Table 10 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Type
Params
dmodel
Layers
Heads
KV
dffn
Experts
top- k
LLaMA-3.1-8B
Dense
8B
4,096
32
32
8
14,336
—
—
LLaMA-3.1-70B
Dense
70B
8,192
80
64
8
28,672
—
—
LLaMA-3.3-70B
Dense
70B
8,192
80
64
8
28,672
—
—
Qwen2.5-72B
Dense
72B
8,192
80
64
8
29,568
—
—
Qwen3.5-9B
Dense
9B
4,096
32
16
4
12,288
—
—
Qwen3.5-27B
Dense
27B
5,120
64
24
4
17,408
—
—
Appendix
Table 8: Model configurations evaluated. Dense models use grouped-query attention (GQA); MoE models use top- k expert routing. Models marked with a dagger ( † ) are included only in the NCU kernel profiling dataset ( kernels_labeled ) and were not run through the serving benchmark.
Model
2080Ti
RTX 3090
A100-40GB
H100
LLaMA-3.1-8B
1, 2, 4
1, 2, 4, 8
1, 2, 4, 8
1, 2, 4
LLaMA-3.1-70B
—
8
4, 8
1, 2, 4
LLaMA-3.3-70B
—
—
4, 8
1, 2, 4
Qwen2.5-72B
—
—
4, 8
1, 2, 4
Mixtral-8 × 7B
—
8
4, 8
2, 4
Qwen3.5-9B
2, 4
1, 2, 4, 8
1, 2, 4, 8
1, 2
Appendix
Table 9: Tensor-parallel configurations actually run in the released serving summaries. Entries list observed TP degrees after collapsing multi-GPU labels within each platform family. gpt-oss models use MXFP4 weights; all others bfloat16.
Profile
OIeff
Ridge multiple
Chat multi-turn
181
0.6×
Computer-use multi-turn
886
3.0×
Terminal multi-turn
1,809
6.1×
Coding-agent multi-turn
2,409
8.2×
Appendix
Table 12: Effective serving OI at C=80 for multi-turn profiles (LLaMA-3.1-8B, H100, TP=1, vLLM, bf16). Values correspond to Figure 4 a.
Profile
ISL
OIprefill
OIdecode
Gap
Chat multi-turn
833
634
78
8×
Computer-use multi-turn
4,666
1,695
78
22×
Terminal multi-turn
10,472
2,123
78
27×
Coding-agent multi-turn
12,929
2,208
78
28×
Appendix
Table 13: Per-phase OI gap at C=80 (LLaMA-3.1-8B, H100, BF16, TP=1). OIprefill is the per-layer GEMM aggregate at L=ISL tokens; OIdecode is the GEMM aggregate at decode batch L=C=80 tokens. The H100 ridge is 295 FLOP/byte; prefill is compute-bound for all profiles, decode is bandwidth-bound.