It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assigned a binary or ordinal score and the aggregation of scores across prompts is then used as a summary evaluation. In this paper, we develop a Bayesian approach for quantifying the uncertainty that arises in such evaluation metrics as a result of the stochasticity of the LLM-based systems; the same prompt may exhibit different outcomes on repeated runs. Our framework leads naturally to a sequential evaluation, in which we leverage the Bayesian model to preferentially select which prompts in the benchmark to use next, enabling more cost-effective LLM evaluations. We demonstrate this approach through four case studies: pairwise LLM preferences in interactive dialogue (MT-Bench), refusal rates on adversarial inputs (JailbreakBench), embodied agentic reasoning (ALFWorld), and agentic web-search Q&A (GAIA). Our results illustrate how a Bayesian approach can provide insightful uncertainty quantification about the behavior of blackbox LLM-based systems in a manner that is not accessible via deterministic or single-run evaluations.
Figures & tables
Figure 1: Posterior distributions of three aggregates given n stochastic generations per prompt, one column per benchmark. Top: W>ν (Poisson binomial). Middle and bottom: Wmean and Wmin (10,000 Monte Carlo draws). In the top two rows, the dashed gray line is the value a single greedy-decoding run reports.
Figure 2: Bayesian approaches (Greedy and Thompson sampling) achieve faster reductions in posterior variance and yield more accurate estimates of E[W>ν] under smaller sampling budgets. We report the sequential algorithms for E[W>ν] , Var(W>ν) and Mean Absolute Error (MAE) of E[W>ν] relative to the posterior given every generation: Appendix C.3 gives two alternative references for MAE. Results are averaged over 1000 runs. The shaded regions indicate the interquartile ranges (25th-75th percentiles).
Figure 3: Kernel density estimates of the posterior distribution of W>ν (Poisson binomial) with a smaller sampling budget. The dashed gray line indicates the conclusion drawn from greedy decoding or a single run. W>ν is discrete; the density estimate is for visualization only.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Prompt used to obtain generations for each prompt in the refusal rate experiments. {{harmful_content}} denotes the content to be replaced with the corresponding prompt. Examples of prompts from JailbreakBench are “defamatory article claiming that a U.S. president is addicted to heroin” and “sexist email about a female colleague”.
Figure 5: Prompt used to obtain generations for each prompt in the pairwise comparison experiments. {{model_content}} denotes the content to be replaced with the corresponding prompt. Examples of prompts from MT-Bench are “compose an engaging travel blog post about a recent trip to Hawaii, highlighting cultural experiences and must-see attractions” and “describe a vivid and unique character, using strong imagery and creative language. Please answer in fewer than two paragraphs”.
Figure 6: Prompt used to obtain evaluations for each prompt in the pairwise preferences experiments. {{question}} denotes the content to be replaced with the corresponding prompt, which is the same as the {{model_content}} shown in Figure 5 . {{answer_a}} and {{answer_b}} denote the content to be replaced with two models’ responses, respectively.
Figure 7: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm using MT-Bench.
Figure 8: Results using the sequential algorithms on JailbreakBench for W>ν with ν=0.95 , M=100 , averaged over 1000 runs.
Figure 9: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm for JailbreakBench.
Figure 10: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm using ALFWorld, at a budget of 8×M .
Figure 11: Average number of times across the 1000 runs that each input prompt is selected by the sequential algorithm for GAIA, at a budget of 8×M .
Figure 12: Posterior distribution of W>ν (Poisson Binomial). Intervals represent 5/95 percentiles. The dashed gray line indicates the conclusion drawn from greedy decoding/one run.
Figure 13: W>ν distributions for M=100 . ϵ=1e−6 , ν=0.95 . Non-informative Jeffreys prior Beta(0.5,0.5) . Averaged over 1000 runs.
Benchmark
Wpost (Figure 2 )
Round-robin endpoint
Wtrue
MT-Bench
1.41
1.27
1.39
JailbreakBench
1.03
1.03
1.03
ALFWorld
3.01
3.01
1.02
GAIA
3.40
3.40
1.11
Appendix
Table 1: Ratio of round-robin’s MAE of E[W>ν] to greedy’s at 10×M under each reference; values above one favor greedy. 1,000 replications.
ν
GAIA
ALFWorld
MT-Bench
JailbreakBench
0.50
60%
40%
30%
84%
0.60
53%
53%
n/a
72%
0.70
47%
27%
52%
70%
0.75
–
–
48% ∗
–
0.80
40%
20%
60%
54%
0.90
40% ∗
47% ∗
n/a
n/a
Appendix
Table 2: Budget saved by greedy against round-robin at the MAE round-robin only reaches at the full reporting budget: 15×M on GAIA and ALFWorld, whose 20 stored generations per prompt are exhausted by every strategy at 20×M , and 50×M on MT-Bench and JailbreakBench. ∗ marks the main-paper ν ; n/a, greedy never reaches and holds that MAE; –, not swept. The last row counts thresholds at which greedy’s MAE at 10×M is below round-robin’s. 1,000 replications per configuration (200 for JailbreakBench).
Budget
GAIA
ALFWorld
MT-Bench
JailbreakBench
0.5×M
29%
35%
50%
14%
1×M
67%
66%
100%
26%
2×M
100%
100%
100%
51%
3×M
100%
100%
100%
76%
4×M
100%
100%
100%
100%
Appendix
Table 3: Share of prompts that greedy has sampled at least once. Round-robin reaches 100% at 1×M by construction. Budgets below 1×M are resolved with checkpoints every M/4 pulls. Jeffreys prior, 1,000 replications.
2×M
5×M
Benchmark
Greedy
Round-robin
Greedy
Round-robin
GAIA
4.24
5.03
2.56
3.26
ALFWorld
3.37
2.69
1.60
1.76
MT-Bench
6.85
9.71
3.04
4.89
JailbreakBench
69.52
65.26
69.95
60.41
Appendix
Table 4: Sensitivity to the prior: the largest minus the smallest MAE of E[W>ν] against Wtrue across the five priors. Smaller is less sensitive. 1,000 replications.
Budget per system
Allocation
Selects better system
Identical allocation
0.5×M
Round-robin
73.5%
100%
Independent greedy
91.8%
6.3%
Joint greedy
84.1%
100%
1×M
Round-robin
93.9%
100%
Independent greedy
99.8%
0.1%
Joint greedy
99.4%
100%
Appendix
Table 5: Comparing two systems with M=50 , ν=0.9 and Δtrue=10 ; 1,000 replications, averaged over three random orderings of the prompts. Identical allocation is the fraction of runs in which both systems received the same number of generations on every prompt.
Figure 14: Sequential evaluation of two LLM judges on M=100 HelpSteer2 (prompt, response) pairs ( Wang et al., 2024b ) . Each generation is one judge call that returns a helpfulness score in {1,…,5} , and the aggregate is W>0.75(5) , the number of items the judge rates 5 with probability above 0.75. Calls are drawn from each judge’s score distribution, read from its next-token probabilities over the five score tokens, so Wtrue is known exactly (dotted line, first column): 10 for Qwen3.6-27B and 19 for Gemma-4-31B. The first two columns show the Dirichlet–multinomial model as in Figure 2 , averaged over 1,000 runs with the interquartile range shaded. The last two columns add the ordinal-probit model (dashed): the MAE of E[W] against Wtrue , and the share of runs whose central 90% credible interval covers Wtrue (dotted line: nominal 0.9).
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environment. Acting rationally then requires inferring the unobserved quantities that govern it and updating beliefs about them as evidence accumulates. Yet most evaluations only score the model's final-turn answer in a single-turn format, leaving this process unexamined. We ask how closely LLMs' belief updates match those of a rational Bayesian reasoner in multi-turn settings, and introduce BayesBench, a suite of simulation environments that probe this across three progressively complex tasks: (i) Bayesian estimation, where the model infers an unknown parameter from sequential evidence; (ii) Bayesian prediction, where the model turns inferred beliefs about a latent variable into outcome forecasts; and (iii) latent-framed Bayesian prediction, where observations are filtered through a user-persona framing, requiring joint inference over the latent state and the persona. Across seven LLMs (3B--70B), scaling improves latent inference and evidence accumulation, with updates occasionally matching the Bayesian posterior. However, these gains do not reliably carry over to downstream prediction, exposing a gap between inferring latent structure and using it to rationally update beliefs about the target outcome.
Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7
Meta AI · Columbia University · Work done at Meta +2
Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example pair to evaluate, thereby avoiding wasted evaluations on clearly underperforming models. Further savings can be achieved by predicting model scores from the partially observed model-example score matrix using low-rank factorization. However, such predictions are not ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap predicted scores without compromising statistical validity. Specifically, we derive doubly robust estimators of each model's performance that use the low-rank predictions to reduce variance. This enables the construction of valid finite-sample confidence intervals in our setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks show that our approach reduces the number of required evaluations, yielding meaningful savings in compute and cost while accurately identifying the best-performing model.
Elad Tolochinsky, Yaniv Tenzer, Yaniv Romano
Department of Computer Science, Technion – Israel Institute of Technology · Department of Electrical and Computer Engineering, Technion – Israel Institute of Technology
We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capability on this question set and to design active querying rules that shrink the CS width as quickly as possible. For CS construction, we invert a family of test supermartingales and focus on two representative approaches: a reverse information projection (RIPr)-based approach and a testing-by-betting-based approach. We first study these approaches under an oracle setting, and demonstrate the oracle optimality of the RIPr-based construction. We then propose a growth-oriented querying rule that aims to maximize the worst-case one-step expected log-increment over the endpoints of the current CS. In practice, we build these test supermartingales and the querying rule on predictions of question-level correctness learned from historical data. We then analyze the shrinkage behavior of the resulting CSs and identify two key factors that slow the shrinkage rate of CSs: accumulated prediction mismatch and the spikiness of the querying distribution. Finally, motivated by this analysis, we propose several mixture querying rules that combine growth-oriented querying, prediction refinement, and uniform exploration, trying to mitigate the effects that slow the shrinkage rate. We provide experiments comparing different querying rules for the RIPr-based and testing-by-betting-based CSs across several synthetic testing datasets. Interestingly, we observe that the simplest querying rule, uniform sampling, can sometimes outperform more adaptive querying rules for both methods.