cs.CLAug 27, 2026

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

Authors: Mingqi Gao, Anthony Sicilia, Weiyan Shi

Organizations: Northeastern University · West Virginia University

Abstract

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased. We develop parametric and non-parametric procedures, analyze the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation. PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non-verifiable tasks.

Explore similar work

CardsList
  1. Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

    Aug 2, 2026Shengwei Xu, Yuxuan Lu, Yifan Wu +2Language Model Generation EvaluationAutomated Evaluation

  2. Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

    Jun 12, 2026Alyssa Unell, Natalie Dullerud, Naomi Boneh +4LLM-as-a-JudgeLanguage Model Generation Evaluation