stat.MLOct 6, 2026

Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects

Authors: Tianxi Li, Jie Ding

Organizations: School of Statistics, University of Minnesota

Abstract

Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

    Oct 4, 2026Laurène Vaugrante, Thilo Hagendorff

  2. Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

    May 11, 2026Elad Tolochinsky, Yaniv Tenzer, Yaniv RomanoLarge Language Model EvaluationModel Selection