Evaluating Rubric Generation with Interventional Transfer
Organizations: Johns Hopkins Computer Science · Amazon
Abstract
Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.
Figures & tables
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric | Description | Location |
|---|---|---|
| RubricRAG [ 6 ] , with scored by an LLM | ||
| LLM-matching | §3.3.1; prompt in Fig. 3 | |
| -gram-matching | , , as above with -gram similarity in place of | §3.3.1 |
| Missed@ , Hallucinations@ , Redundancy@ | §3.3.2 | |
| Score-correlation | , Spearman and Pearson is the human-authored response | §4.3, Table 3 |
| Discrimination between high/low quality responses | for each of is generated by conditioning on another example’s rubric | §3.4; Fig. 8 |
| Metric | Description | Location |
|---|---|---|
| GenRubric [ 3 ] , over responses from different models for each example, with and | ||
| Score-correlation | called “sorting consistency” | App. B.1 |
| Kendall’s | , for concordant and discordant response pairs, and , the pairs tied under , | App. B.2 |
| Top-1 consistency | App. B.3 | |
| Pairwise accuracy | App. B.4 | |
| aggregation | per query, averaged over generation runs, then over queries | App. B.5 |
| Setting | Choice |
|---|---|
| Judge (scores a response against a rubric) | Gemini 3.8 Flash, reasoning effort medium |
| Judge votes | majority of 3 samples |
| Negative criteria restated as positive | Gemini 3.8 Flash, reasoning effort low |
| LLM-matching framings ( Figure A.1 ) | Gemini 3.8 Flash, reasoning effort medium |
| RubricRAG similarity | Gemini 3.8 Flash, reasoning effort low |
| Generators | Qwen3.8-27B, Deepseek-V4-Flash, Opus-5 |
| Deepseek-V4-Flash Qwen3.8-27B | Opus-5 Qwen3.8-27B | Opus-5 Deepseek-V4-Flash | |
|---|---|---|---|
| LLM-matching | 0.04 (0.02, 0.05) | 0.02 (0.01, 0.04) | -0.01 (-0.03, 0.00) |
| Score Correlation | 0.04 (-0.05, 0.13) | 0.06 (-0.02, 0.15) | 0.02 (-0.06, 0.09) |
| Score Correlation (physician) | -0.05 (-0.21, 0.09) | 0.03 (-0.14, 0.19) | 0.07 (-0.10, 0.27) |
| Discrimination between high/low quality responses | 0.05 (-0.03, 0.15) | 0.13 (0.03, 0.23) | 0.08 (-0.02, 0.18) |
| expert score | ||||
|---|---|---|---|---|
| response model | ordinary | foreign rubric | difference | share of passing lost |
| Qwen | 0.58 (0.54, 0.62) | 0.41 (0.37, 0.46) | 0.17 (0.13, 0.21) | 0.27 (0.22, 0.33) |
| Llama-4-Scout-17B | 0.45 (0.41, 0.50) | 0.34 (0.30, 0.38) | 0.11 (0.07, 0.15) | 0.25 (0.17, 0.31) |
| Qwq-32B | 0.58 (0.54, 0.62) | 0.46 (0.42, 0.51) | 0.12 (0.07, 0.16) | 0.21 (0.16, 0.26) |
| Glm-5.3-Flash | 0.61 (0.57, 0.65) | 0.55 (0.50, 0.59) | 0.07 (0.03, 0.11) | 0.09 (0.04, 0.15) |
| Grok-4.20 | 0.58 (0.53, 0.62) | 0.36 (0.32, 0.41) | 0.22 (0.16, 0.27) | 0.37 (0.28, 0.45) |
| rubric generator | passing | failing | passing failing | ||
|---|---|---|---|---|---|
| Qwen3.8-27B | 4.42 | 825 | 3.28 | 134 | 1.14 (0.67, 1.57) |
| Deepseek-V4-Flash | 4.80 | 754 | 3.87 | 174 | 0.92 (0.47, 1.36) |
| Opus-5 | 4.09 | 1269 | 3.43 | 397 | 0.65 (0.31, 1.02) |
| passing / failing | |||
|---|---|---|---|
| similarity | Qwen3.8-27B | Deepseek-V4-Flash | Opus-5 |
| 1 | 95% / 91% | 98% / 98% | 93% / 92% |
| 2 | 87% / 75% | 91% / 86% | 83% / 78% |
| 3 | 75% / 54% | 81% / 69% | 69% / 59% |
| 4 | 58% / 39% | 66% / 48% | 52% / 40% |
| 5 | 46% / 28% | 53% / 34% | 41% / 30% |
| Deepseek-V4-Flash Qwen3.8-27B | Opus-5 Qwen3.8-27B | Opus-5 Deepseek-V4-Flash | |
|---|---|---|---|
| Training target | 0.05 (-0.02, 0.13) | 0.11 (0.04, 0.17) | 0.06 (-0.02, 0.13) |
| Detect good responses | -0.03 (-0.18, 0.09) | -0.09 (-0.22, 0.04) | -0.05 (-0.18, 0.10) |
| Detect bad responses | 0.10 (0.05, 0.14) | -0.00 (-0.04, 0.04) | -0.10 (-0.14, -0.06) |
| Degradation flag | -0.03 (-0.07, 0.01) | 0.03 (0.00, 0.07) | 0.06 (0.03, 0.10) |
| whole rubric, | single criterion, | |||||||
|---|---|---|---|---|---|---|---|---|
| target | measured | |||||||
| , | ||||||||
| , | ||||||||
| , | ||||||||
| , | ||||||||