Three Ways Classical Test Theory Can Mislead About LLM Judges
Organizations: University of Oxford
Abstract
Evaluations that use a large language model (LLM) as a judge have begun to borrow reliability statistics from classical test theory and its extensions. We examine three such statistics that need one administration and no gold labels. None of them can isolate the judge, because one judge under one prompt supplies no variance component of its own. Claude Haiku 4.5 judged 210 constructed short answers against ten-element checklists. On the 180 with parsed verdicts, the Kuder-Richardson coefficient (KR-20) came out at 0.5223 on the judge's verdicts and 0.5231 on error-free gold verdicts. In simulation, bank design alone moves KR-20 from 0.01 to 0.68 at the judge's measured 4.72% error rate. The dependability index , a ratio of mean squared distances from the pass mark, sits 0.22 to 0.38 below the judge's accuracy against gold and returns 0.54 to 0.68 on error-free gold verdicts. Livingston-Lewis accuracy treats the rubric elements as a sample, and at a pass mark of five elements it credits error-free gold scores with 0.78, close to the judge's 0.81. A statement about the judge therefore needs gold labels or a varied scorer facet, and a reliability ratio needs the bank's spread beside it. One of the four closest judge-evaluation papers varies the prompt and still reads a reliability below 0.7 as a sign that a model cannot serve as a judge, although that reliability moves with the spread of the samples scored. We derive a decision table and four reporting lines from these two rules.
Figures & tables
| Question | Quantity that answers it | Requires | Tempting substitute |
|---|---|---|---|
| How often does the judge’s pass or fail match gold? | Accuracy against gold, with cluster intervals | Gold labels | , a ratio of squared distances, or Livingston–Lewis accuracy, which counts element sampling as error |
| Where does the judge err, and in which direction? | False-accept and false-reject rates per checklist line | Gold labels | A whole-rubric reliability coefficient |
| Do the judge’s decisions depend on the wording of its instruction? | Phrasing variance components from a person-by-phrasing study (Appendix A.5 ) | Varied phrasings | Agreement across repeated samples at one prompt, which measures decode-time noise |
| Do the checklist elements rank responses consistently? | KR-20 | Neither | KR-20 on judge verdicts read as a property of the judge |
| Would pass or fail survive another sample of checklist elements? | Livingston–Lewis consistency and accuracy | Neither | Accuracy against gold, which holds the elements fixed |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Artefact | File | Contents |
|---|---|---|
| Bank | judge_item_bank.csv | 210 responses with question id, question and answer text, element text, element-level gold ( ) and judge ( ) verdicts and totals; judge fields blank for the 30 unparsed responses |
| Estimators | estimators.py | KR-20, KR-21, inter-element correlation, beta-binomial MLE and goodness of fit, Livingston–Lewis DC and CA, variance components and , cluster bootstrap |
| Bank statistics | bank_stats.py | Every quantity measured on the real bank, with cluster intervals and the independence check of Section 3 |
| Simulations | sweeps.py | Two-way sweep, measured-rate run, one-directional sweep, control, Livingston–Lewis estimand control, phrasing identification check |
| Published grids | sweep_grid.json | Grid values as reported here, including the 60-replicate run at the measured error rate |
| Figures | make_figures.py | Regenerates Figures 2 , A.1 and A.2 from the released data; Figure 1 is a schematic |
| Question. How does a refrigerator move heat out of its interior? |
| Answer as presented to the judge. A refrigerant fluid circulates through a closed loop of coils. Evaporation releases heat into the interior, which is then vented outside separately. Expansion causes the refrigerant to cool to a temperature below the interior. The cycle repeats continuously, net moving heat from cold interior to warm exterior against the natural gradient. Heat release at the condenser causes the refrigerant to condense into a liquid. The cold refrigerant absorbs heat from the interior at the evaporator coils. |
| Element position | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Gold rate | 0.506 | 0.583 | 0.561 | 0.533 | 0.550 | 0.567 | 0.561 | 0.528 | 0.544 | 0.578 |
| Judge error | 0.144 | 0.044 | 0.072 | 0.061 | 0.044 | 0.033 | 0.033 | 0.000 | 0.022 | 0.017 |
| Judge per-element error | ||||||
|---|---|---|---|---|---|---|
| Bank spread | 0% | 4.72% | 5% | 10% | 20% | 30% |
| 0.0 | -0.000 | 0.007 | -0.014 | -0.011 | 0.015 | 0.003 |
| 0.2 | 0.111 | 0.098 | 0.092 | 0.062 | 0.068 | -0.009 |
| 0.4 | 0.366 | 0.317 | 0.328 | 0.253 | 0.174 | 0.045 |
| 0.6 | 0.575 | 0.522 | 0.519 | 0.451 | 0.325 | 0.168 |
| 0.8 | 0.722 | 0.679 | 0.682 | 0.608 | 0.457 | 0.266 |
| Bank spread | 0% | 4.72% | 5% | 10% | 20% | 30% |
|---|---|---|---|---|---|---|
| 0.0 | 0.087 | 0.085 | 0.111 | 0.106 | 0.091 | 0.112 |
| 0.2 | 0.097 | 0.083 | 0.097 | 0.078 | 0.093 | 0.100 |
| 0.4 | 0.061 | 0.068 | 0.063 | 0.085 | 0.088 | 0.111 |
| 0.6 | 0.044 | 0.047 | 0.048 | 0.041 | 0.065 | 0.084 |
| 0.8 | 0.021 | 0.030 | 0.028 | 0.043 | 0.059 | 0.078 |
| Accuracy | LL CA | LL CA | Accuracy | LL CA | LL DC | ||
|---|---|---|---|---|---|---|---|
| Cut | vs gold | interval | judge | gold | LL CA | (gold) | judge |
| 4 | 0.961 | 0.922–0.994 | 0.882 | 0.846 | 0.079 | 0.154 | 0.836 |
| 5 | 0.928 | 0.893–0.962 | 0.812 | 0.781 | 0.116 | 0.219 | 0.755 |
| 6 | 0.900 | 0.851–0.945 | 0.759 | 0.750 | 0.141 | 0.250 | 0.697 |
| 7 | 0.917 | 0.856–0.973 | 0.750 | 0.769 | 0.167 | 0.231 | 0.689 |
| Scenario | ||||
|---|---|---|---|---|
| Phrasing irrelevant | 0.994 | 0.000 | 0.090 | 0.916 |
| Mild phrasing effect | 0.994 | 0.039 | 0.090 | 0.886 |
| Large phrasing effect | 0.994 | 0.365 | 0.090 | 0.721 |
| Deterministic given phrasing | 0.996 | 0.177 | 0.000 | 0.865 |
| Line | Worked wording |
|---|---|
| Bank | Reliability computed over binary checklist elements on 180 responses from 13 question clusters. SD of gold totals on a – scale. Mean inter-element correlation . |
| Gold comparison | KR-20 is on judge verdicts and on gold verdicts. Livingston–Lewis accuracy at a pass mark of five elements is on judge totals and on gold totals. Because error-free scoring lands in the same region, neither value describes the judge. |
| Dependability coefficient on judge verdicts and on gold verdicts. This is a ratio of squared distances from the pass mark and is not comparable to the classification accuracy of at the same pass mark. | |
| Judge error | Against gold, the judge accepts of absent elements ( cluster interval – ) and rejects of present ones ( – ). Five of the 130 checklist lines hold 25 of the 85 errors. Gate accuracy at a pass mark of five elements is ( – ) on the 180 scored responses and on all 210 when unparsed outputs count as errors. |
| Term | Meaning |
|---|---|
| Response | One constructed short answer, the unit the judge scores; the released file calls it an item |
| Bank | The set of responses a judge scores, here 210 constructed short answers |
| Element | One binary checklist criterion, with per question and a separate checklist for each question |
| Element position | The index of an element within its question’s checklist |
| Gold verdict | Whether element is present in response , fixed by construction |
| Judge verdict | The judge’s YES ( ) or NO ( ) for element of response |