Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
Organizations: Citigroup Inc., London, United Kingdom · Ernst & Young LLP, London, United Kingdom · NVIDIA Corporation, London, United Kingdom
Abstract
LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.
Figures & tables
| Device | Weight tier / capacity | Bandwidth | Peak (dense, stated) | Execution | Status |
|---|---|---|---|---|---|
| NVIDIA H100 / H200 | HBM3/3e, 80 / 141 GB | 3.35 / 4.8 TB/s | 1 PFLOPS FP8 | SIMT | Installed base |
| NVIDIA B200 / B300 | HBM3e, 192 / 288 GB | 8 TB/s | 4.5 PFLOPS FP8, FP4 | SIMT | Volume |
| NVIDIA GB200 NVL72 rack | 72 GPUs; 13.4 TB HBM3e | 130 TB/s NVLink aggregate | rack-scale | single NVLink domain | 120–140 kW liquid rack |
| NVIDIA Rubin | HBM4, 288 GB | 20+ TB/s | 50 PFLOPS FP4 (claim) | SIMT | Ramp H2 2026 |
| Groq 3 LPU / LPX rack (NVIDIA) | on-die SRAM 0.5 GB/LPU; 128 GB/256-LPU rack | 150 TB/s per LPU | n/d | static TSP | Decode co-processor, Q3 2026 |
| Groq LPU v1 (GroqCloud) | on-die SRAM 230 MB | 80 TB/s | 188 TFLOPS FP16 | static TSP | Cloud only |
| Model | Hardware | tok/s/user | Source, class |
|---|---|---|---|
| Llama 3.1 8B / 70B | Cerebras CS-3 | 1,800 / 2,100 | ( Cerebras Systems, 2024 ) D-vendor |
| Llama 3.1 405B | Cerebras CS-3 | 969 | ( Cerebras Systems, 2024 ) D-vendor |
| Llama 4 Scout | Cerebras CS-3 | 2,600 | ( Artificial Analysis, 2026 ) D-indep |
| Kimi K2.6 ( 1T) | Cerebras CS-3 | 1,000 | ( Cerebras Systems, 2026b ; Artificial Analysis, 2026 ) D-indep |
| GPT-OSS-120B | Cerebras CS-4 | 4,400 | ( Cerebras Systems, 2026a ) D-vendor |
| Llama 3 70B | Groq LPU v1 | 284 (TTFT 0.3 s) | ( Groq, 2024 ) D-indep |
| Layer | Metric | Definition / why | |
|---|---|---|---|
| L1 | weight-tier BW, capacity | Eq. 1 – 2 ; STREAM-style kernel | E |
| L1 | tokens/J (batch 1, 8) | wall-plug energy ( Luccioni et al., 2024 ; Samsi et al., 2023 ) | E |
| L1 | determinism | CoV of TPOT over identical runs | P |
| L2 | TTFT, TPOT P50/P99 vs context | MLPerf Interactive definitions | E |
| L2 | goodput | Eq. 3 at the agent SLO | E |
| L2 | prefix-cache hit rate; KV-transfer P99 | context grows each step; disaggregation adds a hop | P |
| Family | Tax. | Stress | Outputs; reference envs |
|---|---|---|---|
| S1 single-step | A1 | 1k–128k in; 64–2k out | TTFT, TPOT, tok/J; MLPerf Interactive |
| S2 long-context RAG | A1–2 | 2k–128k retrieved | TTFT, KV footprint, accuracy ( Liu et al., 2024 ) |
| S3 tool-augmented agent | A2 | 5–30 steps, 1–8 tools | , , tool share; -bench, GAIA |
| S4 long-horizon agent | A5 | 20–200 steps, retries 0–3 | , time-to-success, P99, ; SWE-bench, WebArena, TheAgentCompany |
| S5 concurrent agents | A2 conc. | 1–256 trajectories | goodput, P95/P99, fairness |
| S6 multi-agent | A3 | 2–8 agents; shared vs private KV | completion, , network bytes; ( Wu et al., 2023 ; Cemri et al., 2025 ) |
| Axis | Levels | Metrics | Hypothesis |
|---|---|---|---|
| A1 context | 2k, 8k, 32k, 128k | TTFT, TPOT vs | H1 |
| A2 output length | 16, 128, 1k, 8k | TPOT stability, tok/J | static devices flat |
| A3 concurrency | 1–256 agents | at , 40 ms; P99 | H2 |
| A4 format | text / JSON-schema / code | TPOT ratio, validity | overhead in runtime, not silicon |
| A5 tool interleave | ; ms, 0.3 s, 2 s | , idle fraction, KV-transfer P99 | H3 |
| A6 faults & reasoning budget | injected ; budget 1/2/4 | measured , , CoV, Eq. 6 | H4 |
| Hypothesis | Test / falsifier | |
|---|---|---|
| H1 | On SRAM decode engines TPOT is flat in context until KV exceeds on-die capacity, then step-wise; on HBM GPUs it rises smoothly. | A1, batch 1 / smooth SRAM degradation |
| H2 | At ms P99, SRAM pools reach the goodput per rack of Blackwell-class pools; at ms . | A3 sweep / ratio invariant to |
| H3 | Episode-time gains saturate once ; for – s the value of TPOT ms is negligible. | A5 / continued gains below 2 ms |
| H4 | Deterministic execution yields lower and TPOT CoV, hence lower , on S4/S6. | A6 fault injection / no difference |
| H5 | Disaggregated pairs give lower cost per successful episode than either homogeneous pool for A2–A4, not A1. | Eq. 6 across classes / homogeneous wins |
| H6 | In regulated firms time-to-reproducible-evidence, not peak throughput, predicts adoption (Eq. 8 ). | Retrospective procurement analysis |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Power | Cooling / footprint |
|---|---|---|
| Cerebras CS-3 | 23-25 kW, 15U | liquid; 2 per rack 50 kW |
| Cerebras CS-4 rack | n/d (3 wafers) | liquid; rack-scale |
| DGX H100 / B200 | 10-14 kW and above | air (H100) / liquid (B200) |
| GB200 / GB300 NVL72 | 120-140 kW | liquid, mandatory |
| Groq 3 LPX rack | n/d | liquid, dense |
| SambaRack 16 SN40L | 10-15 kW | air, high airflow |
| Vendor | Status; anchors and channels; access modes; residual risks |
|---|---|
| Cerebras | Status: public (Nasdaq: CBRS, May 2026); 2025 revenue US$510 M; CS-4 announced Aug 2026. Anchors: OpenAI 750 MW; AWS Bedrock decode tier; G42/MBZUAI concentration. Access: appliance; Cerebras Cloud; Bedrock; HPE and Supermicro. Risks: customer concentration; wafer yield; export licences; single-appliance availability. |
| NVIDIA (incl. Groq LPU) | Status: incumbent; US$20 B non-exclusive Groq licence; Groq 3 LPX ships Q3 2026. Anchors: universal. Access: Vera Rubin racks with LPX decode tier; Dynamo disaggregated serving. Risks: licence-structure scrutiny; LPX capacity; premium pricing. |
| Groq (independent) | Status: GroqCloud continues; core team departed. Anchors: developer cloud. Access: API only. Risks: roadmap and talent for on-prem hardware. |
| SambaNova | Status: private; US11 B (Jul 2026). Anchors: JPMorganChase on-prem; SoftBank SN50; US DOE labs. Access: SambaRack; managed co-location; SambaCloud; Xeon 6 heterogeneous design. Risks: SN50 execution; support scale; proprietary stack. |
| Intel | Status: Gaudi 3 de-emphasised; Crescent Island sampling H2 2026. Anchors: IBM Cloud, Dell. Access: OEM. Risks: roadmap credibility. |
| AMD | Status: MI355X shipping; leads MLPerf v6.0 Llama-2-70B rows. Anchors: hyperscalers; OpenAI (reported). Access: OEM; ROCm. Risks: ecosystem maturity. |