It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assigned a binary or ordinal score and the aggregation of scores across prompts is then used as a summary evaluation. In this paper, we develop a Bayesian approach for quantifying the uncertainty that arises in such evaluation metrics as a result of the stochasticity of the LLM-based systems; the same prompt may exhibit different outcomes on repeated runs. Our framework leads naturally to a sequential evaluation, in which we leverage the Bayesian model to preferentially select which prompts in the benchmark to use next, enabling more cost-effective LLM evaluations. We demonstrate this approach through four case studies: pairwise LLM preferences in interactive dialogue (MT-Bench), refusal rates on adversarial inputs (JailbreakBench), embodied agentic reasoning (ALFWorld), and agentic web-search Q&A (GAIA). Our results illustrate how a Bayesian approach can provide insightful uncertainty quantification about the behavior of blackbox LLM-based systems in a manner that is not accessible via deterministic or single-run evaluations.
Figures & tables
Figure 1: Posterior distributions of three aggregates given n stochastic generations per prompt, one column per benchmark. Top: W>ν (Poisson binomial). Middle and bottom: Wmean and Wmin (10,000 Monte Carlo draws). In the top two rows, the dashed gray line is the value a single greedy-decoding run reports.
Figure 2: Bayesian approaches (Greedy and Thompson sampling) achieve faster reductions in posterior variance and yield more accurate estimates of E[W>ν] under smaller sampling budgets. We report the sequential algorithms for E[W>ν] , Var(W>ν) and Mean Absolute Error (MAE) of E[W>ν] relative to the posterior given every generation: Appendix C.3 gives two alternative references for MAE. Results are averaged over 1000 runs. The shaded regions indicate the interquartile ranges (25th-75th percentiles).
Figure 3: Kernel density estimates of the posterior distribution of W>ν (Poisson binomial) with a smaller sampling budget. The dashed gray line indicates the conclusion drawn from greedy decoding or a single run. W>ν is discrete; the density estimate is for visualization only.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Prompt used to obtain generations for each prompt in the refusal rate experiments. {{harmful_content}} denotes the content to be replaced with the corresponding prompt. Examples of prompts from JailbreakBench are “defamatory article claiming that a U.S. president is addicted to heroin” and “sexist email about a female colleague”.
Figure 5: Prompt used to obtain generations for each prompt in the pairwise comparison experiments. {{model_content}} denotes the content to be replaced with the corresponding prompt. Examples of prompts from MT-Bench are “compose an engaging travel blog post about a recent trip to Hawaii, highlighting cultural experiences and must-see attractions” and “describe a vivid and unique character, using strong imagery and creative language. Please answer in fewer than two paragraphs”.
Figure 6: Prompt used to obtain evaluations for each prompt in the pairwise preferences experiments. {{question}} denotes the content to be replaced with the corresponding prompt, which is the same as the {{model_content}} shown in Figure 5 . {{answer_a}} and {{answer_b}} denote the content to be replaced with two models’ responses, respectively.
Figure 7: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm using MT-Bench.
Figure 8: Results using the sequential algorithms on JailbreakBench for W>ν with ν=0.95 , M=100 , averaged over 1000 runs.
Figure 9: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm for JailbreakBench.
Figure 10: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm using ALFWorld, at a budget of 8×M .
Figure 11: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm for GAIA, at a budget of 8×M .
Figure 12: Posterior distribution of W>ν (Poisson Binomial). Intervals represent 5/95 percentiles. The dashed gray line indicates the conclusion drawn from greedy decoding/one run.
Figure 13: W>ν distributions for M=100 . ϵ=1e−6 , ν=0.95 . Non-informative Jeffreys prior Beta(0.5,0.5) . Averaged over 1000 runs.
Benchmark
Wpost (Figure 2 )
Round-robin endpoint
Wtrue
MT-Bench
1.41
1.27
1.39
JailbreakBench
1.03
1.03
1.03
ALFWorld
3.01
3.01
1.02
GAIA
3.40
3.40
1.11
Appendix
Table 1: Ratio of round-robin’s MAE of E[W>ν] to greedy’s at 10×M under each reference; values above one favor greedy. 1,000 replications.
ν
GAIA
ALFWorld
MT-Bench
JailbreakBench
0.50
60%
40%
30%
84%
0.60
53%
53%
n/a
72%
0.70
47%
27%
52%
70%
0.75
–
–
48% ∗
–
0.80
40%
20%
60%
54%
0.90
40% ∗
47% ∗
n/a
n/a
Appendix
Table 2: Budget saved by greedy against round-robin at the MAE round-robin only reaches at the full reporting budget: 15×M on GAIA and ALFWorld, whose 20 stored generations per prompt are exhausted by every strategy at 20×M , and 50×M on MT-Bench and JailbreakBench. ∗ marks the main-paper ν ; n/a, greedy never reaches and holds that MAE; –, not swept. The last row counts thresholds at which greedy’s MAE at 10×M is below round-robin’s. 1,000 replications per configuration (200 for JailbreakBench).
Budget
GAIA
ALFWorld
MT-Bench
JailbreakBench
0.5×M
29%
35%
50%
14%
1×M
67%
66%
100%
26%
2×M
100%
100%
100%
51%
3×M
100%
100%
100%
76%
4×M
100%
100%
100%
100%
Appendix
Table 3: Share of prompts that greedy has sampled at least once. Round-robin reaches 100% at 1×M by construction. Budgets below 1×M are resolved with checkpoints every M/4 pulls. Jeffreys prior, 1,000 replications.
2×M
5×M
Benchmark
Greedy
Round-robin
Greedy
Round-robin
GAIA
4.24
5.03
2.56
3.26
ALFWorld
3.37
2.69
1.60
1.76
MT-Bench
6.85
9.71
3.04
4.89
JailbreakBench
69.52
65.26
69.95
60.41
Appendix
Table 4: Sensitivity to the prior: the largest minus the smallest MAE of E[W>ν] against Wtrue across the five priors. Smaller is less sensitive. 1,000 replications.
Budget per system
Allocation
Selects better system
Identical allocation
0.5×M
Round-robin
73.5%
100%
Independent greedy
91.8%
6.3%
Joint greedy
84.1%
100%
1×M
Round-robin
93.9%
100%
Independent greedy
99.8%
0.1%
Joint greedy
99.4%
100%
Appendix
Table 5: Comparing two systems with M=50 , ν=0.9 and Δtrue=10 ; 1,000 replications, averaged over three random orderings of the prompts. Identical allocation is the fraction of runs in which both systems received the same number of generations on every prompt.
Figure 14: Sequential evaluation of two LLM judges on M=100 HelpSteer2 (prompt, response) pairs ( Wang et al., 2024b ) . Each generation is one judge call that returns a helpfulness score in {1,…,5} , and the aggregate is W>0.75(5) , the number of items the judge rates 5 with probability above 0.75. Calls are drawn from each judge’s score distribution, read from its next-token probabilities over the five score tokens, so Wtrue is known exactly (dotted line, first column): 10 for Qwen3.6-27B and 19 for Gemma-4-31B. The first two columns show the Dirichlet–multinomial model as in Figure 2 , averaged over 1,000 runs with the interquartile range shaded. The last two columns add the ordinal-probit model (dashed): the MAE of E[W] against Wtrue , and the share of runs whose central 90% credible interval covers Wtrue (dotted line: nominal 0.9).
Department of Computer Science, Technion – Israel Institute of Technology · Department of Electrical and Computer Engineering, Technion – Israel Institute of Technology