cs.CLNov 4, 2025

Sequential Bayesian Evaluation of Large Language Model Behavior

Authors: Saatvik Kher, Shang Wu, Rachel Longjohn, Catarina Belém, Padhraic Smyth

Organizations: Department of Computer Science, University of California, Irvine · Department of Statistics, University of California, Irvine

Abstract

It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assigned a binary or ordinal score and the aggregation of scores across prompts is then used as a summary evaluation. In this paper, we develop a Bayesian approach for quantifying the uncertainty that arises in such evaluation metrics as a result of the stochasticity of the LLM-based systems; the same prompt may exhibit different outcomes on repeated runs. Our framework leads naturally to a sequential evaluation, in which we leverage the Bayesian model to preferentially select which prompts in the benchmark to use next, enabling more cost-effective LLM evaluations. We demonstrate this approach through four case studies: pairwise LLM preferences in interactive dialogue (MT-Bench), refusal rates on adversarial inputs (JailbreakBench), embodied agentic reasoning (ALFWorld), and agentic web-search Q&A (GAIA). Our results illustrate how a Bayesian approach can provide insightful uncertainty quantification about the behavior of blackbox LLM-based systems in a manner that is not accessible via deterministic or single-run evaluations.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

    Jun 29, 2026Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7LLM Reasoning Strategies

  2. Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

    May 11, 2026Elad Tolochinsky, Yaniv Tenzer, Yaniv RomanoLarge Language Model EvaluationModel Selection

  3. Efficient Sequential Evaluation of Large Language Models

    Jul 19, 2026Chia-Yu Hsu, Shubhanshu ShekharLarge Language Model EvaluationConfidence Calibration