cs.CLAug 27, 2026

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

Authors: Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang

Organizations: The University of Chicago · Johns Hopkins University · The Hong Kong Polytechnic University, Hong Kong

Abstract

LLM-judge audits assess bias by comparing ratings across matched conditions. Difference-in-differences designs compare two candidate responses within each item and then compare that contrast across a manipulated attribute. We show in closed form that this endpoint need not identify differential preference on the latent scale when ratings are bounded. A severity shift common to both responses produces an observed interaction whenever the scale bounds attenuate it unequally. We examine this problem in a pre-registered audit comprising 990 calls to a frozen pedagogy judge. The registered primary analysis did not detect an effect of the stated learner profile on scaffolding preference. The only nominally significant secondary interaction concerned productive struggle (+0.378+0.378 points; p=0.002p = 0.002). A post-hoc construction with zero differential preference reproduced 79 to 85% of this interaction using the observed severity shift and the scale floor alone. The observed interaction therefore does not identify differential preference on the latent scale.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix