Rubrics distill notions of expert quality and measure agent performance. However, the quality of rubrics themselves have not been systematically measured and are often left to downstream performance.We import apparatuses from measurement theory built for exactly this: quantitative signals based on the rubric's content, and introduce the RubrIc-Failure Taxonomy (RIFT), of nine possible ways a rubric fails, organized under reliability and content validity. Every mode leaves a distinct signature. To show the signals track failure causally, we seed 720 corruptions, injecting each RIFT mode into clean rubrics at known severity levels. A linear probe over the signals identifies which mode was injected at 75.0% accuracy, beating 56.7% for a frontier model asked to name the failure directly. Surprisingly across GDPval and Terminal-Bench, 10 of 48 expert-authored rubrics weight their criteria backwards, putting more of the score on requirements an expert panel judged less essential. This means a response can fail what matters most and still be graded well. This paper serves as a comprehensive guide on how to understand failure modes in rubrics and create better versions using quality signals, causal experiments, and provides a taxonomy with its rules and examples.
Figures & tables
Figure 1: This paper grounds rubric design in well-established quality signals. These help us understand and view causal effects on rubric design. We also provide a taxonomy (RIFT) to help designers name specific failure modes.
Figure 2: The RubrIc Failure Taxonomy (RIFT) organizes nine modes of rubric failure. See table 2 for complete definitions.
Figure 3: Failure-mode signatures at ρ=0.5 : paired Cohen’s d (variant − baseline) for each of 8 injected modes across 17 signals, n=60 tasks per mode. Asterisks mark cells significant under Benjamini–Hochberg FDR at q<.05 applied across the full 8×17 grid (46 of 136). Blank cells are signals whose delta was identically zero on every task.
Figure 4: Left: distribution of ρord over the 60 human-authored rubrics ( n=48 scoreable); red bars are anti-calibrated. Right: the three weight-calibration signals under weight-shuffle injection, mean ±95% CI over 35 paired tasks against mean realized ρ .
Human
Clean
ρ=.25
ρ=.50
ρ=.75
ρord
+0.195
+0.189
−0.522
−0.562
−0.562
Core weight share
0.835
0.717
0.471
0.246
0.083
FSR
0.950
0.956
0.974
0.988
0.996
Table 1: Weight-calibration signals on human rubrics ( n=60 , panel-scored) and under weight-shuffle injection ( n=35 paired). Higher ρord and core weight share are better; higher FSR is worse.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Sub-dimension
Failure mode
Reliability
Ambiguity
Subjective Language : evaluative terms left unanchored to objective expectations
Ungrounded Verification : checkable facts with no reference, source, or procedure
Missing Criteria : a prompt requirement that no criterion evaluates
Over-Specification
Rigid : stricter or narrower requirements than the prompt asks for
Redundant Criteria : one behavior scored by two or more criteria
Appendix
Table 2: RIFT: RubrIc Failure Mode Taxonomy. Nine failure modes under two quality dimensions. Full definitions, decision rules, and pass/fail examples in Appendix B .
Mode
GPT-5.5
3-judge panel
p
Hackable
−2.83
−2.51
2.6×10−27
Rigid
−2.04
−1.89
2.9×10−21
Misaligned
−1.85
−2.24
8.5×10−25
Appendix
Table 3: Paired Cohen’s d for the mean CVR collapse (variant − clean, n=60 /mode) under GPT-5.5 versus an independent 3-judge panel ( claude-opus-4-8 , claude-haiku-4-5 , gemini-3.1-pro ).
AUC .13 , p=5×10−30 ; loudest effect in the corpus
weighted_relevance
−2.07
item_relevance
−1.94
candidate_item_mean
−1.94
core_pct
−1.17
AUC .29
Appendix
Table 4: Per-mode signatures at ρ=0.5 ( n=60 per mode). Paired Cohen’s d (variant − baseline); all entries shown are FDR-significant at q<.05 over the full 8×17 grid. k is the number of significant signals for that mode. AUC is the unpaired Mann–Whitney P(variant>baseline) , where 0.5 indicates a signal blind to the corruption.
Figure 5: Signal response against realized corruption fraction ρ ( 95% CI, n=15 paired tasks per level). The CVR-based signals and total_items descend steadily with corruption, while redundancy_score and ungrounded verification’s weighted relevance stay flat across the full range.
Mode
Signal
Slope
r
p
GDPval / TB
Graded: slope estimates severity
Hackable
Δ mean_cvr
−1.55
−.83
1.6×10−12
−1.83 / −1.21
Rigid
Δ mean_cvr
−1.35
−.73
1.6×10−8
−1.62 / −1.01
Misaligned
Δ mean_cvr
−1.10
5×10−7
Missing criteria
Δ total_items
−13.3
<.001
Mechanical
Non-atomic
Δ weighted_relevance
+0.159
+.67
<.001
Inverted
Appendix
Table 5: OLS slope of Δ signal on realized ρ over the 15-task subset (3 corruption levels, 45 paired points per mode). Per-family slopes give the same regression fit within the GDPval and Terminal-Bench subsets independently. Realized corruption fractions average 0.267 , 0.528 , and 0.766 at the three requested levels, with non-overlapping ranges.
GDPval: EMS cardiac quality-assurance review. A paramedic reconstructs county EMS STEMI cases for a medical-direction QA report.
EMS medical director
Focused on whether STEMI recognition, alert activation, and protocol adherence are clinically justified by the available field data.
Cardiac QA paramedic
A senior QA reviewer assessing reconstructed EMS timelines, ECG timing, documentation consistency, and field decision-making.
Interventional cardiologist
Evaluates whether prehospital STEMI suspicion aligns with ECG findings, ED assessment, and cath lab outcomes.
Protocol compliance officer
Determines whether ECG acquisition timing and STEMI alert decisions meet county EMS standards.
Documentation auditor
Checks that conclusions are supported by reconciled evidence from PCRs, monitor exports, hospital notes, and protocols.
GDPval: B2B SaaS release retrospective. A QA lead analyses defect leakage after Release 4.7 and recommends test-plan changes.
Appendix
Table 6: Generated persona panels for four tasks. Each panel is produced once per task from the instruction alone and reused across all experimental conditions. Profile lines are verbatim; task descriptions are our one-line summaries.
Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over adversarial answers optimized to exploit them. Yet their robustness to such optimization remains poorly understood. We isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is to acknowledge the impossibility. We introduce ImpossibleRubrics, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls. Rather than providing fixed rubrics, ImpossibleRubrics provides task environments and certificates, allowing rubrics to be generated downstream and then adversarially tested for whether they reward certificate-violating answers. Eleven generators are exploited 8--26% of the time on the unbiased 150-of-169 environment cut; on a deliberately selected stress cut the strongest generator we measured is still exploited 36% while a certificate-faithful rubric is exploited 0%, so what we measure is a rubric-quality gap, not task impossibility. One result runs against intuition. A single generic rubric ("be decisive, penalize hedging") used unchanged for every task is exploited 64% of the time, and seven of the eleven generators are exploited more often than that while writing a rubric tailored to each one. The tailored criteria appear to tell an attacker which claim to fabricate. The problem is not that rubrics are vague; it is that they are specific about the wrong things.
Bowen Qin, Yi Xie, Yesheng Liu +1
National University of Singapore · Peking University · Institute of Automation, Chinese Academy of Sciences
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.
Milad Yazdani, Yaser Souri, Xiren Zhou +4
Department of Electrical and Computer Engineering, University of British Columbia · Microsoft · School of Biomedical Engineering, University of British Columbia
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric
Maoqi Liu, Junwei He, Bowen Zhang +5
Beijing University of Posts and Telecommunications · ByteDance