Rater State Bias in RLHF Preference Data: An Audit Framework
Authors: Elena Kopteva, Vitaliy Hlynianyi-Zhuk
Organizations: The Grainger College of Engineering, Department of Physics & Illinois Center for Advanced Studies of the Universe, University of Illinois Urbana-Champaign, Urbana, Illinois 61801, USA · Faculty of Applied Mathematics, Oles Honchar Dnipro National University, Dnipro, 49045, Ukraine · Department of Clinical Psychology, Kyiv Institute of Modern Psychology and Psychotherapy, Kyiv, 01133, Ukraine
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time, so that preference data encode rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from ordinary disagreement or random label noise. They would be state dependent, could be shared across annotators under similar conditions, and would not necessarily cancel during aggregation, reward modeling, and policy optimization. We propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also propose survival level emotional authenticity as a candidate output signature, defined by lexical, pragmatic, discourse, and safety features whose reliability and validity remain to be demonstrated. We show that systematic rater state bias can survive aggregation and may enter the learned reward signal. We state five testable predictions, together with effect size thresholds for an initial audit, and note which require proprietary data. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model.
Figures & tables
f
δ
ρ
Deff
Neff
f⋅δ
0.10
0.10
0.00
1.0
1000
0.010
0.10
0.10
0.10
20.9
48
0.010
0.20
0.15
0.00
1.0
1000
0.030
0.20
0.15
0.10
20.9
48
0.030
0.20
0.15
0.20
40.8
25
0.030
Table 1: Illustrative estimates of rater state bias entry. f : fraction of annotations made under shifted conditions. δ : preference probability shift under those conditions. ρ : intraclass correlation. Deff : design effect ( n=200 ). Neff : effective independent sample size ( N=1,000 ). The mean shift fδ is invariant to ρ ; the effective sample size is not.
Stage
Mechanism
Effect on bias
Preference collection
Rater state shift (Eq. 5 )
The aggregate preference signal shifts by fδ(x) . Positive correlation reduces Neff , increasing uncertainty in estimates of the shift.
Reward model training
Bradley-Terry fitting (Eq. 7 )
A structured shift in preference probabilities changes the fitted reward differences between responses.
Policy optimization
Reward weighting with a KL penalty (Eq. 9 )
An absorbed reward shift changes relative response odds by exp(Δr/λ) .
Table 2: Propagation of rater state bias through RLHF. Each stage can preserve or amplify the shift introduced at the previous stage.
Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape large language model behaviour. However, the normative role of these judgments is rarely made explicit. I distinguish three conceptual models of that role. The first is extension: annotators extend the system designers' own judgments about what outputs should be. The second is evidence: annotators provide independent evidence about some facts, whether moral, social or otherwise. The third is authority: annotators have some independent authority (as representatives of the broader population) to determine system outputs. I argue that these models have implications for how RLHF pipelines should solicit, validate and aggregate annotations. I survey landmark papers in the literature on RLHF and related methods to illustrate how they implicitly draw on these models, describe failure modes that come from unintentionally or intentionally conflating them, and offer normative criteria for choosing among them. My central recommendation is that RLHF pipeline designers should decompose annotation into separable dimensions and tailor each pipeline to the model most appropriate for that dimension, rather than seeking a single unified pipeline.
Steve Coyne
University of Toronto, Canada · University of Toronto, Toronto, Canada
How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outputs are used to train a reward model that assigns scalar values to responses. Because these rewards are inferred from pairwise comparisons, this learning depends on an assumed relationship between latent reward differences and observed preferences, typically modeled using a Boltzmann formulation in which a rationality parameter beta informs how consistently preferences reflect reward differences. In practice, beta is typically treated as a fixed constant that reflects assumed uniform annotator reliability. However, human feedback is not this simplistic in practice: real human judgments are shaped by cognitive biases, leading to systematic deviations from reward-consistent behavior that arise contextually. To address this, we treat rationality as context- and annotation-dependent. We design an approach to dynamically adjust the rationality parameter beta during reward learning using an LLM-as-judge to assess the likely presence of cognitive biases. This approach effectively downweights comparisons that are likely to reflect biased or unreliable judgments. Empirically, we show that this approach learns a more rational downstream model, even when finetuning on datasets with strongly biased preferences.
Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to amplify undesired behaviors. This arises from core limitations of RLHF: (1) preference datasets are constructed from the LLM's own outputs, allowing it to influence them, and (2) pairwise comparisons only indicate which response is better, not why. These limitations can be exploited to cause alignment tampering. For example, if an LLM generates biased responses with higher quality, annotators will prefer them based on quality. However, preference labels do not distinguish quality from bias, and the reward model inherits this limitation. Optimizing such rewards through reinforcement learning or best-of-N sampling can amplify misaligned biases. Our experiments demonstrate amplification across diverse biases: from keyword bias to propaganda (e.g., sexism), brand promotion, and instrumental goal-seeking. Mitigation remains challenging, as existing techniques for robust RLHF fail to fully resolve alignment tampering without sacrificing response quality. These findings reveal structural vulnerabilities of current RLHF and emphasize the need to prevent this vulnerability. Project page: https://alignment-tampering.github.io/