Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
Organizations: The University of Chicago · Johns Hopkins University · The Hong Kong Polytechnic University, Hong Kong
Abstract
LLM-judge audits assess bias by comparing ratings across matched conditions. Difference-in-differences designs compare two candidate responses within each item and then compare that contrast across a manipulated attribute. We show in closed form that this endpoint need not identify differential preference on the latent scale when ratings are bounded. A severity shift common to both responses produces an observed interaction whenever the scale bounds attenuate it unequally. We examine this problem in a pre-registered audit comprising 990 calls to a frozen pedagogy judge. The registered primary analysis did not detect an effect of the stated learner profile on scaffolding preference. The only nominally significant secondary interaction concerned productive struggle ( points; ). A post-hoc construction with zero differential preference reproduced 79 to 85% of this interaction using the observed severity shift and the scale floor alone. The observed interaction therefore does not identify differential preference on the latent scale.
Figures & tables
| Weak | Strong | |||
|---|---|---|---|---|
| Arm | 95% BCa | 95% BCa | ||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value | Breakdown |
| Stimuli | 55 | 30 weak / 25 strong |
| Source tutoring runs | 23 | 18 weak / 10 strong / 5 shared |
| Word problems, all acid-mixture | 6 | |
| Single-message contexts, no tutor turn | 23 | 7/30 weak, 16/25 strong |
| Stimuli by source-run tutor type | 31 / 24 | ped. / conv., from 13 / 10 runs |
| Real turn is the high pole | 43/55 | 29/31 ped., 14/24 conv. |
| Quantity | Role in the paper | Location |
| Registered; computed by the frozen, hashed analysis code | ||
| Preference table (arm stratum) | Baseline scaffolding preference, both strata | Table 1 |
| Primary endpoint | The registered result; null | § 6.1 |
| Strong-stratum companion | Secondary; agrees with the primary | § 6.1 |
| Pure profile effect | Shows the profile reached the rating | § 6.2 |
| Evidence-gradient check | Tests the anchoring reading; null | § 6.2 |