Certified Selective Automation of LLM Agent Evaluation
Authors: Chengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni
Organizations: Independent Researcher · University of Chinese Academy of Sciences · Pusan National University · Shenzhen University of Advanced Technology
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.
Figures & tables
Figure 1: Certified selective automation, on one corpus. (a) A calibrated threshold auto-rejects low-scoring trajectories; the certificate bounds the error rate among auto-decided trajectories . (b) Attempts at one task are correlated: resample tasks, not trajectories.
Figure 2: Overview. One certificate fixes the deployment operating point (§ 4 ), is predictable from corpus structure before training (§ 6 ), and gates a self-training loop whose contamination is bounded by the same α (§ 7 ).
corpus
domain
n
G
π
ρ
DEFF
neff
Tool ( τ2 )
tool-use dialogues
1280
84
.48
.25
5.0
257
Term
terminal sessions
520
13
.34
.43
17.6
30
Web-A (ARB)
live web, expert labels
200
68
.26
.49
2.2
91
Web-M
MiniWoB++, 4B agents
512
32
.16
.80
13.0
39
Web-F
MiniWoB++, frontier agents
256
32
.26
.81
6.7
38
Code-O
code repair (OpenHands)
899
382
.45
.64
2.0
459
Table 1: Corpora. n : test trajectories; G : test tasks; π : success rate; ρ : intra-task correlation of outcomes; DEFF=1+(m~−1)ρ with m~ the size-weighted mean cluster size; neff=n/DEFF .
Figure 3: TaskBoot in one picture: a pre-declared quantile grid (a), a task-resampled test of each threshold’s selective error (b), max-coverage selection with calibrate-on-CAL / audit-on-TEST deployment (c).
Figure 4: Certifiability is measurable ex-ante. Best certified reject coverage against the pre-training index (1−π)(2A0−1) . The largest residual ( Term ) is the smallest- neff corpus, where finite-sample slack in the certificate binds before judge quality does.
AUROC
certified coverage @ α=0.1
corpus
base
SFT
GRPO
base
SFT
GRPO
ΔRL
regime
Tool
.798
.899
.903
.000
.293
.297
+.004
trainable
Web-A
.897
.925
.922
.510
.560
.585
+.025
trainable
Web-M
.983
.976
.977
.832
.789
.836
+.047
saturated
Web-F
.983
.991
.994
.707
.758
.758
±.000
saturated
Code-O
.531
.645
.530
.000
.000
.000
—
judge-blind
Table 2: Main result: certified reject coverage at α=0.1 , δ=0.05 ( TaskBoot , threshold calibrated on CAL), and test AUROC. GRPO arm selected on CAL. Grey: G below the simulation-validated regime ( G≥20 ), certificate reported for completeness only. Oracle-arm and α=0.2 results in Appendix E .
judge
elicitation
Web-A
Tool
AUROC
cert@.1
AUROC
cert@.1
gpt-5.6-sol
CoT + prob.
.929
.704
.794
.000
claude-sonnet-5
CoT + prob.
.905
.497
.847
.000
gemini-2.5-pro
CoT + prob.
.753
.000
—
—
gemini-2.5-pro
SC-5
.680
.000
—
—
claude-sonnet-5
SC-10
.686
.000
—
—
Table 3: Strongly elicited frontier judges vs. the trained 4B judge, certified by the same procedure. CoT: chain-of-thought then a verbalized probability; SC- k : vote share over k samples at T=1 ; logprob: native token probability. Cost: relative inference cost per decision.
target corpus
judge
cert@.1
harvest (contam.)
target labels
Web-A
transfer (round 0)
.575
—
0
+ CertHarvest round 1
.585
262/914 (.008)
0
175 CAL labels spent on SFT instead
.540
—
175
in-domain ceiling (SFT → GRPO, § 5 )
.585
—
914
Web-M
transfer (round 0)
.830
—
0
+ CertHarvest round 1
.820
837/1296 (.002)
0
Table 4: Certificate-gated self-training (certified reject coverage @ α=0.1 ); transfer judges never saw the target corpus.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
G
ρ
naive CP
DEFF-CP
one-per-task
task-Hoeffding
TaskBoot
20
.1
.00 / .65
.00 / .16
.00 / .00
.00 / .00
.00 / .68
20
.5
.00 / .53
.00 / .00
.00 / .00
.00 / .00
.01 / .61
20
.8
.00 / .47
.00 / .00
.00 / .00
.00 / .00
.01 / .56
50
.1
.00 / .69
.00 / .67
.00 / .00
.00 / .00
.00 / .69
50
.5
.00 / .61
.00 / .03
.00 / .00
.00 / .00
.00 / .62
50
.8
.00 / .57
.00 / .00
.00 / .00
.00 / .00
.00 / .57
Appendix
Table 5: Full synthetic grid: violation / mean certified coverage.
corpus
judge
G
ρ
naive
DEFF
one/task
Hoeffding
TaskBoot
Tool
base
84
.25
.000
.000
.000
.000
.000
Tool
SFT
84
.25
.325
.000
.000
.000
.325
Tool
GRPO μ3
84
.25
.321
.000
.000
.000
.321
Term
base
13
.43
.340
.000
.000
.000
.315
Term
SFT
13
.43
.379
.000
.000
.000
.379
Term
GRPO μ5
13
.43
.369
.000
.000
.000
.369
Appendix
Table 6: Certified reject coverage @ α=0.1 on real test scores, five procedures. One-per-task uses a single pre-registered draw (seed 42); task-Hoeffding tests the mean of per-task error rates over covered tasks (estimand: task-weighted error), Bonferroni δ/40 throughout.
corpus
SFT
acc. ( λ=μ=1 )
reject μ=3
reject μ=5
release λ=5
CAL pick
Tool
.293
.297 (.150)
.321 (.149)
.295 (.145)
.293
μ=1
Web-A
.560
.585 (.514)
.585 (.537)
.585 (.514)
.585
μ=3
Appendix
Table 11
target
start
harvest
contamination
TEST cert @ .1
Web-A
transfer, round 1 (full-CAL pilot)
437/914
.041
.585
Web-A
transfer, round 1 (split-CAL)
262/914
.008
.585
Web-A
transfer, round 2
336/914
.024
.465
Tool
in-domain, round 1
296/2200
.000
.343
Tool
in-domain, round 2
338/2200
.003
.347
Web-M
transfer, round 1
837/1296
.002
.820
Appendix
Table 7: All CertHarvest harvests. The full-CAL pilot row shows the protocol before the split-CAL fix; its result is unchanged by the fix, and all reported numbers use the split-CAL protocol.
form
in-sample R2
LOOCV R2
LOOCV MAE
(1−π)(2A0−1) (Eq. 6 )
.977
.964
.039
(1−π)(2A0−1)(1−e−neff/K) , K=30
.835
.707
.124
same, K tuned by LOOCV ( K=10 )
.977
.967
.039
(2A0−1) only
.913
.471
.132
(2A0−1)(1−e−neff/30) (no π )
.656
.382
.196
Appendix
Table 8: Certifiability-model validation: in-sample and leave-one-corpus-out fit of Eq. ( 6 ) and ablated forms.
Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures whether this gate's ranking matches a grounded verifiable reward across 25 agents from six providers on the τ2-bench and SimulatorArena benchmarks, separating two kinds of evaluation validity that release practices conflate: ranking validity and construct validity. First, a satisfaction-success gap: satisfaction carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer's task, a pattern consistent across five rater populations, both benchmarks, and every subjective dimension we rated. Second, while the gate's ranking is robust across the broad capability span, it loses resolution among the near-equal strong agents: this decision-disagreement rate jumps from <1% on wide-reward pairs to 31% on close pairs. The gate is thus human-validated yet mis-anchored. As a remedy, we propose a calibrate-then-trust cadence in which a judge-free completion bit is a zero-cost tripwire for truncation regressions.
Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, but this assumption has rarely been validated against human annotation. We introduce AgentProp-Bench, a 2,000-task benchmark with 2,300 traces across four domains, nine production LLMs, and a 100-label human-validated subset. We quantify judge reliability, characterize error propagation, and evaluate a runtime mitigation. Substring-based judging agrees with human annotation at kappa=0.049 (chance-level); a three-LLM ensemble reaches kappa=0.432 (moderate) with a conservative bias. Under validated evaluation, a parameter-level injection propagates to a wrong final answer with human-calibrated probability approximately 0.62 (range 0.46-0.73 across models). Rejection (catching bad parameters) and recovery (correcting after acceptance) are independent model capabilities (Spearman rho=0.126, p=0.747). A tuned runtime interceptor reduces hallucination on GPT-4o-mini by 23.0 percentage points under a concurrent n=600 control, but shows no significant effect on Gemini-2.0-Flash, whose aggressive parameter rejection eliminates the target failure mode. All code, data, traces, and human labels are released at https://github.com/bhaskargurram-ai/agenthallu-bench.
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agentic benchmarks can take hours, making it difficult to scale evaluations for measuring performance or curating training data. This has driven widespread reliance on automated approaches such as LLM-as-a-judge (LLMJ) to critique agents at the process and outcome-levels at scale, however, the soundness of LLMJ critiques often goes unmeasured. Here, we introduce Counsel, the first public dataset of meta-evaluations for agentic tasks. Counsel consists of process-level critiques from open-weight LLMJs on two agent benchmarks: tau-bench (customer support agents) and DA-Code (coding agents), and human meta-evaluations of these critiques. Human annotators label critiques on each flagged error as "spot on", "correct location but poor reasoning", or "should not have flagged", achieving reliable inter-annotator agreement (Krippendorff's alpha of 0.78). The resulting dataset stratifies LLMJ critiques by human alignment across both error location within a trajectory and reasoning quality, serving as valuable data to calibrate, improve, or train LLMJs for agents. Comparing open-weight judges, we find that more capable judge models and more reasoning effort both enabled improved human agreement, with the strongest judge reaching ~88% agreement on location and ~65% on reasoning. Counsel is generated using open-weight models and is permissively licensed for broad community use, which we hope will enable rigorous study and improved alignment of LLM-based evaluators for agentic systems.
Sashank Pisupati, Henry Broomfield, Eujeong Choi +5