cs.PFOct 7, 2026

Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing

Authors: Arnold Olympio, Juan Manuel Servera Bondroit, Wael Abdelmalek, Guang Lu, João Carvalho

Organizations: Lucerne University of Applied Sciences and Arts, Switzerland · Microsoft, Switzerland · Uthereal AG, Switzerland · ETH Zurich, Switzerland

Abstract

Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.

Figures & tables

Explore similar work

CardsList
  1. The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility

    May 19, 2026David Pape, Jonathan Evertz, Lea SchönherrLarge Language Model BenchmarksLLM Inference Optimization

  2. What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend

    Aug 5, 2026Shahed Masoudian, Passant Shafaei, Monorama Swain +1Large Language Model BenchmarksLLM Inference Optimization