PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
Organizations: Uniphore
Abstract
Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.
Figures & tables
| retry (r) | draws | kept | rejected | rejected | retention rate (%) | cumulative retention rate (%) |
|---|---|---|---|---|---|---|
| 1,125 | 770 | 266 | 89 | 68.4 | 68.4 | |
| 355 | 161 | 146 | 48 | 45.4 | 82.8 | |
| 193 | 69 | 100 | 24 | 35.8 | 88.9 | |
| total | 1,673 | 1,000 | 512 | 161 | 59.8 | 88.9 |
| dataset | yield | synthetic label accuracy | ||
|---|---|---|---|---|
| (unfiltered) | 150 | 100% | 73.3% [66–80] | +0.468 |
| ( ) | 114 | 76% | 84.2% [76–90] | +0.686 |
| ( ) | 104 | 69% | 84.6% [76–90] | +0.694 |
| parameters | accuracy (%) | leniency | average | tie | by domain (%) | by criterion (%) | # distinct | ||||||||||
| # | evaluator | released | total | active | reas. | mean SD | (mean score) | score SD | (%) | airl. | bank. | retail | telec. | clar. | friend. | task | scores |
| 1 | claude-opus-5 | 2026-07 | closed | closed | Yes | 85.1 0.93 | 0.324 | 0.023 | 4.1 | 84 | 82 | 86 | 87 | 73 | 96 | 85 | 69 |
| 3 | kimi-k3 | 2026-07 | 2.8T | 104B | Yes | 83.6 0.12 | 0.508 | 0.036 | 6.0 | 83 | 80 | 84 | 87 | 71 | 97 | 81 | 37 |
| 6 | glm-5p2 | 2026-06 | 753B | 30B | Yes | 81.5 0.50 | 0.521 | 0.051 | 10.1 | 84 | 81 | 81 | 82 | 78 | 92 | 73 | 31 |
| 8 | gpt-5.4-nano ‡ | 2026-03 | closed | closed | No | 79.4 1.16 | 0.631 | 0.052 | 7.8 | 81 | 78 | 78 | 81 | 74 | 84 | 80 | 61 |
| 10 | gpt-5.6-sol ‡ | 2026-07 | closed | closed | Yes | 78.7 0.40 | 0.543 | 0.035 | 4.9 | 79 | 77 | 81 | 77 | 64 | 97 | 74 | 86 |
| factor | |||
|---|---|---|---|
| tie rate | 25 | ||
| distinct score values | 25 | ||
| leniency (mean score) | 25 | ||
| total parameters | 15 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Generation | |
|---|---|
| a task from the dataset, carrying its domain policy, tools, and user goal. | |
| a criterion, supplied at runtime as free text: a name and a detailed description. | |
| a steer level, , ordered by intended quality. | |
| everything held fixed while the steer level varies: agent model, user simulator, domain, decoding seed. | |
| the steering instruction, written by a small model from , , and , Eq. ( 1 ). | |
| a trajectory: the complete trace generated by an interactive dialogue between a user simulator and an agent executing a task, with tool calls and its outcome. | |
| level | pairs | description |
| Evaluation criterion | ||
| friendliness | 355 | Warmth and consideration toward the customer: whether the agent acknowledges their situation and how they feel about it, delivers unwelcome news with care, and leaves them feeling attended to. Judge the manner, not whether the request was resolved. |
| task resolution | 329 | Whether the customer’s actual problem was settled: did the agent establish what was needed, take the actions that would resolve it, and leave the customer with the outcome they came for. A correct refusal counts as resolution – if the request was not permitted, saying so plainly and explaining why resolves it, while quietly doing it anyway does not. Judge the outcome, not the manner or how well it was explained. |
| communication clarity | 316 | How easily the customer can follow the agent: whether the main point is findable, whether technical or policy language is explained, whether multi-part information is organised, and whether the customer is left knowing what is true and what happens next. Judge the presentation, not the warmth or the outcome. |
| Domain | ||
| retail | 309 | 114 tasks; 1,158-word policy, 16 agent tools |
| filter judge | human annotator | evaluator under test | |
|---|---|---|---|
| purpose | curates the dataset | validates the dataset | the object of measurement |
| is given | the criterion name and definition, the agent’s tools and the domain policy, and both trajectories | the same five fields, rendered in a purpose-built interface (Appendix D ) | the criterion name and definition, the agent’s tools and the domain policy, and one trajectory |
| is not given | the steering prompts, the target levels, or which side was steered better | the same, and no filter verdict | the same, and never the other trajectory |
| is asked | which of the two is better on the criterion | which of the two is better on the criterion | to score the agent on the criterion from to |
| slot order | randomized per pair and per judge | randomized per pair | no order arises: one trajectory per call |
| returns | a forced choice of one side, with at most three sentences of reasoning | a forced choice of one side, with a confidence rating of | a score in , with at most three sentences of reasoning |
| axis | subset | cells filled | retention rate | cum. retention rate | draws per retention |
|---|---|---|---|---|---|
| By criterion | |||||
| friendliness | 355/375 | 70.3% | 95% | 1.42 | |
| task resolution | 329/375 | 58.9% | 88% | 1.70 | |
| communication clarity | 316/375 | 51.9% | 84% | 1.93 | |
| By domain | |||||
| airline | 135/150 | 61.4% | 90% | 1.63 | |
| axis | slot-2 share range | binomial range |
|---|---|---|
| criterion | 48.7–53.8% | 0.17–0.91 |
| domain | 46.9–54.1% | 0.27–0.73 |
| agent model | 45.5–54.3% | 0.11–0.22 |
| level contrast | 49.2–54.0% | 0.14–0.82 |
| whole dataset | 511/1,000 = 51.1% |
| criterion | pairs | mean reward gap | better / worse | sign test |
|---|---|---|---|---|
| task resolution | 329 | 49 / 9 | ||
| communication clarity | 316 | 40 / 14 | ||
| friendliness | 355 | 29 / 23 | 0.488 |
| rater pool | |||
|---|---|---|---|
| , mean SD | |||
| measure | pooled (150) | panel 1 (75) | panel 2 (75) |
| Krippendorff’s | [+0.24, +0.45] | [+0.29, +0.60] | [+0.09, +0.38] |
| unanimous (3–0) | 77/150 = 51% | 44/75 = 59% | 33/75 = 44% |
| Agreement between every two annotators of a panel | |||
| agreement | Cohen’s | both said “certain” | |
| panel 1, | 81.3% (61/75) | 85% (28/33) | |
| panel 1, | 65.3% (49/75) | 75% (24/32) | |
| subset | mean confidence | SD | |
|---|---|---|---|
| (unfiltered) | 150 | 1.49 | 0.43 |
| ( ) | 114 | 1.49 | 0.44 |
| ( ) | 104 | 1.51 | 0.44 |
| Kept against removed, per filter judge | |||
| judge | kept | removed | difference ( ) |
| 1.49 ( , 114) | 1.49 ( , 36) | (1.000) | |
| subset | as reported | domain-weighted | criterion-weighted | agent-weighted | contrast-weighted |
|---|---|---|---|---|---|
| (unfiltered) | 73.3% | 73.1% | 73.3% | 73.4% | 73.7% |
| ( ) | 84.2% | 83.4% | 83.8% | 84.6% | 84.6% |
| ( ) | 84.6% | 83.9% | 84.2% | 85.0% | 85.3% |
| axis | subset | unfiltered ( ) | 1 filter ( ) | 2 filters ( ) | gain | retention |
| By criterion | ||||||
| friendliness | 90% (45/50) | 93% (42/45) | 93% (39/42) | pt | 84% | |
| task resolution | 68% (34/50) | 77% (30/39) | 78% (28/36) | pt | 72% | |
| communication clarity | 62% (31/50) | 80% (24/30) | 81% (21/26) | pt | 52% | |
| By domain | ||||||
| airline | 75% (27/36) | 92% (23/25) | 92% (22/24) | pt | 67% | |
| domain | policy words | agent tools | user tools | sims | median msgs | median tool calls | tool-error sims | mean reward |
|---|---|---|---|---|---|---|---|---|
| airline | 1,313 | 14 | 0 | 440 | 18 | 5 | 21% | 0.434 |
| banking | 926 | 16 | 2 | 855 | 30 | 7 | 3% | 0.037 |
| retail | 1,158 | 16 | 0 | 980 | 22 | 6 | 34% | 0.353 |
| telecom | 3,715 | 13 | 30 | 1,072 | 48 | 10 | 48% | 0.243 |
| human panel | pairs removed | label was right | purity | yield |
|---|---|---|---|---|
| 1 | 6 | 4 | 86.0 88.2% ( pt) | pt |
| 2 | 4 | 4 | 82.5 81.1% ( pt) | pt |
| pooled | 10 | 8 | 84.2 84.6% ( pt) | pt |
| accuracy (%) | |||
|---|---|---|---|
| evaluator | its own pairs | the other pairs | gap (pt) |
| gpt-oss-120b | 86.5 ( ) | 77.4 ( ) | |
| gpt-oss-20b | 68.3 ( ) | 68.9 ( ) | |
| nemotron-lightning-3.5 | 52.1 ( ) | 67.8 ( ) | |
| Pipeline Stage | API Calls | Tokens | Token Share | Uncached Cost | Effective Cost |
|---|---|---|---|---|---|
| 1. Steering Generation | 1,125 | 2.56M | 0.3% | $0.85 | $0.67 |
| 2. Trajectory Generation | 77,980 | 709.4M | 95.0% | $51.16 | $18.36 |
| 3. Filter Judges ( ) | 2,879 | 34.8M | 4.7% | $4.42 | $2.27 |
| Total Pipeline | 81,984 | 746.7M | 100.0% | $56.43 | $23.63 |
| model | prompt tokens | cached | rate |
|---|---|---|---|
| nemotron-lightning-3.5 | 350,084,755 | 312,862,256 | 89.4% |
| gpt-oss-20b | 137,504,779 | 125,776,651 | 91.5% |
| qwen3-30b-a3b-instruct | 107,519,458 | 102,523,608 | 95.4% |
| gpt-oss-120b | 98,571,871 | 86,652,136 | 87.9% |
| total | 693,680,863 | 627,814,651 | 90.5% |
| asset | role in this work | license | terms of use |
|---|---|---|---|
| Benchmark substrate and serving infrastructure | |||
| -bench, sierra-research/ tau2-bench [ 20 , 21 , 22 ] | Task definitions, domain environments, tool sets, agent framework, and user simulator | MIT License | Free commercial and non-commercial use, modification, and redistribution provided the MIT notice (Copyright (c) 2025 Sierra Research) is preserved; evaluation runs are additionally subject to third-party LLM API terms |
| Fireworks AI serverless inference | Hosted inference for every open-weight model in the pipeline and the sweep | not applicable (service) | Fireworks AI terms of service |
| Open-weight models | |||
| gpt-oss-20b , gpt-oss-120b | Trajectory generation; 120b also generates steering instructions and serves as filter judge ; both are evaluators | Apache 2.0 | model license and Fireworks AI terms of service |
| nemotron-lightning-3.5 | Filter judge , trajectory generation, and evaluator | NVIDIA OpenMDW-1.1 | as above |