LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.
Figures & tables
Figure 1: Prefill and decode have opposite bottlenecks. (a) Prefill is compute-bound; decode re-reads all weights per token and is memory-bound. (b) Roofline ( Williams et al., 2009 ) for three memory systems (FP8 dense peaks, indicative): at low-batch decode intensities (red) attainable throughput is set entirely by bandwidth; peak compute matters only at prefill/large-batch intensities (green). Device parameters are vendor specifications [D] .
Figure 2: Reliability compounds along a trajectory [P] . (a) Expected work multiplier (Eq. 5 ) versus trajectory length for p=0.1 – 2% , with and without per-step checkpointing. (b) Expected episode time for a 20-step agent versus TPOT and p : a faster, less reliable engine can lose to a slower, reliable one. Illustrative.
Figure 3: Architectural taxonomy [P] (positions qualitative). The upper-right region is decode-specialised: extreme bandwidth, limited capacity, near-full utilisation at batch 1-a natural partner for an HBM-rich prefill engine rather than a replacement.
Device
Weight tier / capacity
Bandwidth
Peak (dense, stated)
Execution
Status
NVIDIA H100 / H200
HBM3/3e, 80 / 141 GB
3.35 / 4.8 TB/s
≈ 1 PFLOPS FP8
SIMT
Installed base
NVIDIA B200 / B300
HBM3e, 192 / 288 GB
≈ 8 TB/s
≈ 4.5 PFLOPS FP8, FP4
SIMT
Volume
NVIDIA GB200 NVL72 rack
72 GPUs; 13.4 TB HBM3e
130 TB/s NVLink aggregate
rack-scale
single NVLink domain
≈ 120–140 kW liquid rack
NVIDIA Rubin
HBM4, 288 GB
≈ 20+ TB/s
∼ 50 PFLOPS FP4 (claim)
SIMT
Ramp H2 2026
Groq 3 LPU / LPX rack (NVIDIA)
on-die SRAM 0.5 GB/LPU; 128 GB/256-LPU rack
≈ 150 TB/s per LPU
n/d
static TSP
Decode co-processor, Q3 2026
Groq LPU v1 (GroqCloud)
on-die SRAM 230 MB
80 TB/s
188 TFLOPS FP16
static TSP
Cloud only
Table 1: Public specifications of 2026 inference accelerators [D] (vendor-stated; SRAM bandwidths are aggregate on-die and not directly comparable to HBM). Sources: ( NVIDIA, 2026b ; NVIDIA, 2026a ; Groq, 2025 ; Intel, 2025 ; Cerebras Systems, 2024 ; Cerebras Systems, 2026a ; Prabhakar et al., 2024 ; SambaNova Systems, 2026 ; Amazon Web Services and Cerebras Systems, 2026 ; MLCommons, 2026 ) .
Figure 4: Capacity–bandwidth landscape [D] (public specifications; log–log). No SRAM device holds a 70B FP8 model on one chip except the wafer-scale WSE-3; others shard across hundreds of chips or pair with a capacity tier. The 2–3 orders of bandwidth separating the clusters is the source of the decode-speed gap in Fig. 6 .
Figure 5: Disaggregated prefill/decode serving : established in research ( Zhong et al., 2024 ; Patel et al., 2024 ) ; adopted in 2026 by AWS+Cerebras ( Amazon Web Services and Cerebras Systems, 2026 ) , NVIDIA Rubin+Groq 3 LPX ( NVIDIA, 2026b ) and Intel+SambaNova ( SambaNova Systems, 2026 ) .
Figure 6: Documented single-user decode throughput, 2024–2026 [D] (log scale; sources in Table 2 ). Not measured by the authors; not a common protocol.
Figure 7: The latency–throughput frontier is intrinsic to batching [E] (bandwidth model of Eq. 1 , 70B FP8; illustrative). GPUs reach high aggregate throughput only by moving right, raising TPOT; an SRAM pool sits below the MLPerf Interactive limit ( MLCommons, 2026 ) at all batch sizes. A single-stream comparison that does not fix the operating point is under-specified.
Layer
Metric
Definition / why
L1
weight-tier BW, capacity
Eq. 1 – 2 ; STREAM-style kernel
E
L1
tokens/J (batch 1, 8)
wall-plug energy ( Luccioni et al., 2024 ; Samsi et al., 2023 )
E
L1
determinism
CoV of TPOT over 103 identical runs
P
L2
TTFT, TPOT P50/P99 vs context
MLPerf Interactive definitions
E
L2
goodput G(θ1,θ2)
Eq. 3 at the agent SLO
E
L2
prefix-cache hit rate; KV-transfer P99
context grows each step; disaggregation adds a hop
P
Table 3: Metric set by layer (E established, P proposed).
Family
Tax.
Stress
Outputs; reference envs
S1 single-step
A1
1k–128k in; 64–2k out
TTFT, TPOT, tok/J; MLPerf Interactive
S2 long-context RAG
A1–2
2k–128k retrieved
TTFT, KV footprint, accuracy ( Liu et al., 2024 )
S3 tool-augmented agent
A2
5–30 steps, 1–8 tools
Tep , S , tool share; τ -bench, GAIA
S4 long-horizon agent
A5
20–200 steps, retries 0–3
S , time-to-success, P99, p ; SWE-bench, WebArena, TheAgentCompany
S5 concurrent agents
A2 × conc.
1–256 trajectories
goodput, P95/P99, fairness
S6 multi-agent
A3
2–8 agents; shared vs private KV
completion, ω , network bytes; ( Wu et al., 2023 ; Cemri et al., 2025 )
Table 5: Sensitivity axes [P] . Baseline: ReAct loop ( Yao et al., 2023 ) over a fixed tool set, run identically on every system.
Figure 8: TCO model [P] (Eq. 9 ; illustrative). (a) Relative cost per token vs utilisation for a GPU node and an appliance with 10× CapEx and 15× decode goodput at the agentic SLO. (b) Break-even: advantage = capacity ratio / capital ratio. 2024 public price points ( ≈ US2–3MperCS−3vsUS0.3–0.6 M per 8-GPU server) imply a capital ratio near 5 – 10× ; the required capacity ratio is then a measurable threshold.
Hypothesis
Test / falsifier
H1
On SRAM decode engines TPOT is flat in context until KV exceeds on-die capacity, then step-wise; on HBM GPUs it rises smoothly.
A1, batch 1 / smooth SRAM degradation
H2
At θ2≤10 ms P99, SRAM pools reach ≥5× the goodput per rack of Blackwell-class pools; at θ2≥100 ms <2× .
A3 sweep / ratio invariant to θ2
H3
Episode-time gains saturate once nTPOT<τtool ; for τ≈0.3 – 2 s the value of TPOT <2 ms is negligible.
A5 / continued gains below 2 ms
H4
Deterministic execution yields lower p and TPOT CoV, hence lower E[W]/K , on S4/S6.
A6 fault injection / no difference
H5
Disaggregated pairs give lower cost per successful episode than either homogeneous pool for A2–A4, not A1.
Eq. 6 across classes / homogeneous wins
H6
In regulated firms time-to-reproducible-evidence, not peak throughput, predicts adoption (Eq. 8 ).
Retrospective procurement analysis
Table 6: Hypotheses and the experiments that test them [H] .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
System
Power
Cooling / footprint
Cerebras CS-3
≈ 23-25 kW, 15U
liquid; 2 per rack ≈ 50 kW
Cerebras CS-4 rack
n/d (3 wafers)
liquid; rack-scale
DGX H100 / B200
≈ 10-14 kW and above
air (H100) / liquid (B200)
GB200 / GB300 NVL72
≈ 120-140 kW
liquid, mandatory
Groq 3 LPX rack
n/d
liquid, dense
SambaRack 16 × SN40L
∼ 10-15 kW
air, high airflow
Appendix
Table 7: Facility characteristics (vendor-stated or widely reported).
Vendor
Status; anchors and channels; access modes; residual risks
Cerebras
Status: public (Nasdaq: CBRS, May 2026); 2025 revenue US$510 M; CS-4 announced Aug 2026. Anchors: OpenAI 750 MW; AWS Bedrock decode tier; G42/MBZUAI concentration. Access: appliance; Cerebras Cloud; Bedrock; HPE and Supermicro. Risks: customer concentration; wafer yield; export licences; single-appliance availability.
NVIDIA (incl. Groq LPU)
Status: incumbent; US$20 B non-exclusive Groq licence; Groq 3 LPX ships Q3 2026. Anchors: universal. Access: Vera Rubin racks with LPX decode tier; Dynamo disaggregated serving. Risks: licence-structure scrutiny; LPX capacity; premium pricing.
Groq (independent)
Status: GroqCloud continues; core team departed. Anchors: developer cloud. Access: API only. Risks: roadmap and talent for on-prem hardware.
SambaNova
Status: private; US1BSeriesFatUS11 B (Jul 2026). Anchors: JPMorganChase on-prem; SoftBank SN50; US DOE labs. Access: SambaRack; managed co-location; SambaCloud; Xeon 6 heterogeneous design. Risks: SN50 execution; support scale; proprietary stack.
Intel
Status: Gaudi 3 de-emphasised; Crescent Island sampling H2 2026. Anchors: IBM Cloud, Dell. Access: OEM. Risks: roadmap credibility.
As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a central systems challenge. While GPUs dominate current deployments, a growing number of AI accelerators claim advantages for LLM inference, yet it remains unclear under which conditions such accelerators outperform GPUs in practice. Recent inference systems decompose execution into Prefill and Decode phases, which exhibit distinct computational characteristics and latency metrics, commonly captured by time to first token (TTFT) and time per output token (TPOT). This paper presents a phase-aware evaluation of LLM inference performance across GPUs and emerging AI accelerators using a common model, Llama2-7B. By separately measuring Prefill and Decode performance, we reveal that accelerator advantages differ by phase and metric. Our results show that GPUs consistently excel in the compute-intensive Prefill phase, while GroqRack achieves significantly lower TPOT during Decode (batching not currently supported). However, GPUs regain an advantage in Decode throughput as batch size increases. These findings demonstrate that each platform exhibits distinct phase-dependent strengths. We further analyze heterogeneous Prefill/Decode disaggregation across different accelerator platforms, identifying performance gains and the workload and network conditions under which such gains are realized.
Shun Usami, Venkatram Vishwanath, E. Wes Bethel
Department of Computer Science San Francisco State University San Francisco, CA, 94132 · Argonne National Laboratory Lemont, IL, 60439 · Lawrence Berkeley National Laboratory Berkeley, CA, 94720
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks primarily focus on simple single-turn chatbot workloads. LLM applications are increasingly agentic: coding agents, terminal execution systems, and tool-use agents issue multi-turn requests with growing context lengths. We introduce AgentPerfBench, a benchmark suite for agentic inference. It uses real traces from agentic benchmarks, such as SWE-Bench and TerminalBench, alongside standard chat baselines. This enables benchmarking of models on multi-turn tasks involving tool calling, skill utilization, and increasing context lengths. AgentPerfBench also samples from empirical distributions of input length, output length, and turn count derived from the real traces, generating representative synthetic profiles for cheap and accurate measurements on new hardware. In addition, we further find that several existing benchmarks fail to accurately reflect real hardware performance for two key reasons: 1) they do not account for realistic context-length growth, and 2) they measure inference performance without operating at hardware saturation. We discuss these issues in detail and provide rich kernel-level Nsight Compute (NCU) traces to construct a new multi-dimensional roofline model that captures hardware-system limitations in both memory bandwidth and memory capacity footprint. The benchmarking suite then includes automated scripts to identify potential bottleneck conditions on emerging hardware when evaluated with diverse agentic traces. Together, these contributions quantify the chat-to-agentic gap in current inference benchmarks and characterise per-kernel GPU resource utilisation via roofline analysis.
Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin +6
Imperial College London · University of Cambridge · University of Oxford
Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by conventional computational modules. This paper analyzes the fundamental tradeoffs between latency, reliability, and cost in LLM-enabled agentic workflows. We introduce performance models for both LLM and non-LLM agents that capture the relationship between computational effort and output quality, incorporating the impact of reasoning and output tokens for LLM agents using a parametric exponential reliability function. Then, we study the design of sequential workflows under latency and cost constraints. Main results include a water-filling token allocation policy and characterizations of optimal workflow reliability in terms of shadow prices.