cs.LGOct 7, 2026

Evaluating Rubric Generation with Interventional Transfer

Authors: Erik Skalnes, Layne C. Price, Raviteja Anantha, Michael Oberst

Organizations: Johns Hopkins Computer Science · Amazon

Abstract

Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

    May 28, 2026Zijie Wang, Eduardo BlancoLLM-as-a-JudgeLanguage Model Generation Evaluation

  2. GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation

    Aug 30, 2026Yifan Chen, Haitao Li, Qingyao Ai +4LLM-as-a-JudgeLanguage Model Generation Evaluation

  3. Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

    Jul 16, 2026Haocheng Yang, Licheng Pan, Yuan Lu +7Rubric-Based Evaluation