cs.CLSep 28, 2026

PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

Authors: Cheng Chang, Yining Mao, Peng Qi

Organizations: Uniphore

Abstract

Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Counsel: A Meta-Evaluation Dataset for Agentic Tasks

    Jun 19, 2026Sashank Pisupati, Henry Broomfield, Eujeong Choi +5Agentic BenchmarksLlm-As-A-Judge

  2. Preference-Aware Rubric Learning for Personalized Evaluation

    May 29, 2026Yilun Qiu, Xiaoyan Zhao, Yang Zhang +7Large Language Model PersonalizationLarge Language Model Evaluation