Rubrics distill notions of expert quality and measure agent performance. However, the quality of rubrics themselves have not been systematically measured and are often left to downstream performance.We import apparatuses from measurement theory built for exactly this: quantitative signals based on the rubric's content, and introduce the RubrIc-Failure Taxonomy (RIFT), of nine possible ways a rubric fails, organized under reliability and content validity. Every mode leaves a distinct signature. To show the signals track failure causally, we seed 720 corruptions, injecting each RIFT mode into clean rubrics at known severity levels. A linear probe over the signals identifies which mode was injected at 75.0% accuracy, beating 56.7% for a frontier model asked to name the failure directly. Surprisingly across GDPval and Terminal-Bench, 10 of 48 expert-authored rubrics weight their criteria backwards, putting more of the score on requirements an expert panel judged less essential. This means a response can fail what matters most and still be graded well. This paper serves as a comprehensive guide on how to understand failure modes in rubrics and create better versions using quality signals, causal experiments, and provides a taxonomy with its rules and examples.
Figures & tables
Figure 1: This paper grounds rubric design in well-established quality signals. These help us understand and view causal effects on rubric design. We also provide a taxonomy (RIFT) to help designers name specific failure modes.
Figure 2: The RubrIc Failure Taxonomy (RIFT) organizes nine modes of rubric failure. See table 2 for complete definitions.
Figure 3: Failure-mode signatures at ρ=0.5 : paired Cohen’s d (variant − baseline) for each of 8 injected modes across 17 signals, n=60 tasks per mode. Asterisks mark cells significant under Benjamini–Hochberg FDR at q<.05 applied across the full 8×17 grid (46 of 136). Blank cells are signals whose delta was identically zero on every task.
Figure 4: Left: distribution of ρord over the 60 human-authored rubrics ( n=48 scoreable); red bars are anti-calibrated. Right: the three weight-calibration signals under weight-shuffle injection, mean ±95% CI over 35 paired tasks against mean realized ρ .
Human
Clean
ρ=.25
ρ=.50
ρ=.75
ρord
+0.195
+0.189
−0.522
−0.562
−0.562
Core weight share
0.835
0.717
0.471
0.246
0.083
FSR
0.950
0.956
0.974
0.988
0.996
Table 1: Weight-calibration signals on human rubrics ( n=60 , panel-scored) and under weight-shuffle injection ( n=35 paired). Higher ρord and core weight share are better; higher FSR is worse.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Sub-dimension
Failure mode
Reliability
Ambiguity
Subjective Language : evaluative terms left unanchored to objective expectations
Ungrounded Verification : checkable facts with no reference, source, or procedure
Missing Criteria : a prompt requirement that no criterion evaluates
Over-Specification
Rigid : stricter or narrower requirements than the prompt asks for
Redundant Criteria : one behavior scored by two or more criteria
Appendix
Table 2: RIFT: RubrIc Failure Mode Taxonomy. Nine failure modes under two quality dimensions. Full definitions, decision rules, and pass/fail examples in Appendix B .
Mode
GPT-5.5
3-judge panel
p
Hackable
−2.83
−2.51
2.6×10−27
Rigid
−2.04
−1.89
2.9×10−21
Misaligned
−1.85
−2.24
8.5×10−25
Appendix
Table 3: Paired Cohen’s d for the mean CVR collapse (variant − clean, n=60 /mode) under GPT-5.5 versus an independent 3-judge panel ( claude-opus-4-8 , claude-haiku-4-5 , gemini-3.1-pro ).
AUC .13 , p=5×10−30 ; loudest effect in the corpus
weighted_relevance
−2.07
item_relevance
−1.94
candidate_item_mean
−1.94
core_pct
−1.17
AUC .29
Appendix
Table 4: Per-mode signatures at ρ=0.5 ( n=60 per mode). Paired Cohen’s d (variant − baseline); all entries shown are FDR-significant at q<.05 over the full 8×17 grid. k is the number of significant signals for that mode. AUC is the unpaired Mann–Whitney P(variant>baseline) , where 0.5 indicates a signal blind to the corruption.
Figure 5: Signal response against realized corruption fraction ρ ( 95% CI, n=15 paired tasks per level). The CVR-based signals and total_items descend steadily with corruption, while redundancy_score and ungrounded verification’s weighted relevance stay flat across the full range.
Mode
Signal
Slope
r
p
GDPval / TB
Graded: slope estimates severity
Hackable
Δ mean_cvr
−1.55
−.83
1.6×10−12
−1.83 / −1.21
Rigid
Δ mean_cvr
−1.35
−.73
1.6×10−8
−1.62 / −1.01
Misaligned
Δ mean_cvr
−1.10
5×10−7
Missing criteria
Δ total_items
−13.3
<.001
Mechanical
Non-atomic
Δ weighted_relevance
+0.159
+.67
<.001
Inverted
Appendix
Table 5: OLS slope of Δ signal on realized ρ over the 15-task subset (3 corruption levels, 45 paired points per mode). Per-family slopes give the same regression fit within the GDPval and Terminal-Bench subsets independently. Realized corruption fractions average 0.267 , 0.528 , and 0.766 at the three requested levels, with non-overlapping ranges.
GDPval: EMS cardiac quality-assurance review. A paramedic reconstructs county EMS STEMI cases for a medical-direction QA report.
EMS medical director
Focused on whether STEMI recognition, alert activation, and protocol adherence are clinically justified by the available field data.
Cardiac QA paramedic
A senior QA reviewer assessing reconstructed EMS timelines, ECG timing, documentation consistency, and field decision-making.
Interventional cardiologist
Evaluates whether prehospital STEMI suspicion aligns with ECG findings, ED assessment, and cath lab outcomes.
Protocol compliance officer
Determines whether ECG acquisition timing and STEMI alert decisions meet county EMS standards.
Documentation auditor
Checks that conclusions are supported by reconciled evidence from PCRs, monitor exports, hospital notes, and protocols.
GDPval: B2B SaaS release retrospective. A QA lead analyses defect leakage after Release 4.7 and recommends test-plan changes.
Appendix
Table 6: Generated persona panels for four tasks. Each panel is produced once per task from the instruction alone and reused across all experimental conditions. Profile lines are verbatim; task descriptions are our one-line summaries.
Department of Electrical and Computer Engineering, University of British Columbia · Microsoft · School of Biomedical Engineering, University of British Columbia