Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores while violating human intent. Existing mitigations rely on static defenses that cannot adapt to novel exploitation strategies. We propose Adversarial Reward Auditing (ARA), a framework that reconceptualizes reward hacking as a dynamic, competitive game. ARA operates in two stages: first, a Hacker policy discovers reward model vulnerabilities while an Auditor learns to detect exploitation from latent representations; second, Auditor-Guided RLHF (AG-RLHF) gates reward signals to penalize detected hacking, transforming reward hacking from an unobservable failure into a measurable, controllable signal. Experiments across three hacking scenarios demonstrate that ARA achieves the best alignment-utility tradeoff among all baselines: reducing sycophancy to near-SFT levels while improving helpfulness, decreasing verbosity while achieving the highest ROUGE-L, and suppressing code gaming while improving Pass@1. Beyond single-domain evaluation, we show that reward hacking, detection, and mitigation all generalize across domains -- a Hacker trained on code gaming exhibits increased sycophancy despite no reward for this behavior, and an Auditor trained on one domain effectively suppresses exploitation in others, enabling efficient multi-domain defense with a single model.
Figures & tables
Figure 1: ARA framework. A Hacker exploits a frozen reward model while an Auditor gates RL reward.
Scenario
Model
Metric
SFT
PPO
PPO+KL
ODIN
InFoRM
RM Ens.
Adv-RM
RM-FT †
ARA
Sycophancy
Llama-2-7B
Rate ↓
36.2
71.8
57.6
52.3
48.1
51.9
51.4
46.8
39.1
Help ↑
41.3
77.4
67.8
64.1
62.6
65.0
70.6
68.3
76.8
Win ↑
–
60.7
61.4
62.8
61.9
62.1
64.2
63.1
69.4
Llama-3.1-8B
Rate ↓
34.1
67.6
53.8
48.4
43.7
49.8
47.6
43.1
36.4
Help ↑
44.8
79.2
70.1
67.2
65.2
66.9
72.8
70.4
78.6
Win ↑
–
62.3
62.8
64.1
63.7
64.1
66.2
64.8
71.8
Table 1: Main results across three reward-hacking scenarios and three model families. For each scenario we report one alignment metric ( ↓ ), one utility metric ( ↑ ), and GPT-4o win rate against SFT ( ↑ ). Bold indicates the best result among RLHF methods; underlining indicates the runner-up.
Figure 2: Proxy reward versus true task performance during RLHF on Llama-3.1-8B.
Figure 3: Proxy reward Rθ(x,y) vs. alignment with the preference direction ux for Llama-3.1-8B. Genuine responses and verified exploits achieve comparable proxy rewards (horizontal overlap) but separate along ux (vertical axis).
Variant
ux ?
AUC ↑
Syco. ↓
Len. ↓
Code ↓
Random neg.
No
0.58
49.2
238
39.6
Direct ux probe
Yes
0.64
48.6
229
37.4
Shuffled ux in S
Corrupt
0.66
44.1
214
34.0
C -only
Yes
0.70
41.0
181
25.8
No filter
No
0.74
40.3
186
27.9
G -only
No
0.78
38.1
169
21.7
Table 2: Hard-negative construction ablation. AUC is averaged across the three scenarios.
Scenario
Direct ux
C -only
G -only
Full ARA
Syco.
0.41
0.55
0.74
0.79
Len.
0.43
0.57
0.72
0.77
Code
0.46
0.60
0.76
0.81
Avg.
0.43
0.57
0.74
0.79
Table 3: ux -adversarial detection. AUC when the usual association between ux -alignment and exploitation is reversed.
Detection Method
Gen.
Expl.
Δ
Proxy Reward ( Rθ )
0.91
0.89
0.02
GPT-4
0.78
0.34
0.44
Auditor Confidence
0.89
0.21
0.68
Table 4 : Genuine vs. exploit detection. Mean scores for 500 genuine and 500 exploit responses ( Rθ>0.85 ).
Hacker Attack (%)
Auditor AUC
ARA Defense
Train
Syco.
Len.
Code
Syco.
Len.
Code
Syco. ↓
Len. ↓
Code ↓
Ref.
36.2 s
8.1 s
4.2 s
–
–
–
72.4 p
347 p
61.3 p
Syco.
72.4
35.2
28.3
0.82
0.77
0.58
38.4
178
48.2
Length
41.8
52.3
22.6
0.75
0.79
0.55
42.1
162
51.6
Code
58.7
31.4
61.3
0.62
0.59
0.84
54.8
246
19.6
All
68.9
78.1
59.7
–
–
–
39.2
165
21.4
Table 5: Cross-domain attack, detection, and defense. Left: Hacker attack rate ( s SFT baseline). Middle: Auditor AUC. Right: ARA defense metrics ( p PPO baseline).
Figure 4: Gating severity sensitivity analysis. Hacking metrics (red, lower is better) and utility metrics (blue, higher is better).
Figure 5: Gating impact on genuine vs. exploitative responses (Llama-3.1-8B, Rθ>0.85 ). Top: Auditor score distributions. Bottom: proxy vs. gated reward.
Error (%)
Retention (%)
Rank ρ
FPR
FNR
Genuine
Exploit
Genuine
Exploit
Syco. ( γ =2)
6.8
8.6
81.8
6.1
.58
− .04
Len. ( γ =2)
4.4
6.8
78.8
5.4
.72
− .15
Code ( γ =3)
5.6
5.2
80.5
7.3
.75
− .25
Table 6: Gating impact at γ∗ (Llama-3.1-8B). FPR/FNR: misclassification rates. Retention: % of proxy reward preserved after gating. ρ : Spearman correlation between Rθ and Rgated .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Hyperparameter
Value
Hacker (PPO)
Learning rate
1×10−6
Batch size
128
PPO epochs
4
Clip ratio ϵ
0.2
Evasion weight λA
0.5
KL penalty βH
0.1
Appendix
Table 7: Hyperparameters.
Figure 6: Auditor size ablation across all scenarios. Top row: Detection AUC. Bottom row: Mitigation performance (red = hacking metric ↓ , blue = utility metric ↑ ). Gold markers indicate optimal configuration for each task. Different tasks require different capacities: length bias saturates at 25M, sycophancy at 35M, and code gaming at 50M, reflecting increasing exploitation complexity.
Hyperparameter
Value
Syco ↓
Help ↑
Length ↓
ROUGE-L ↑
Code ↓
Pass@1 ↑
Avg. AUC ↑
Default ARA
–
36.4
78.6
157
25.0
18.4
38.4
0.82
λA
0.1
43.0
76.5
190
24.2
29.6
36.8
0.75
λA
0.3
38.9
78.4
165
24.8
21.7
38.1
0.80
λA
0.5
36.4
78.6
157
25.0
18.4
38.4
0.82
λA
0.7
35.8
76.9
154
24.7
17.9
37.2
0.83
λA
1.0
34.9
73.9
151
23.9
17.5
35.8
0.82
Appendix
Table 8: Stage-1 hyperparameter sensitivity. The default is not the best value for every individual metric, but it provides the strongest overall alignment–utility tradeoff. More aggressive Hacker pressure can lower hacking slightly, while usually reducing utility.
Auditor objective
Syco AUC ↑
Length AUC ↑
Code AUC ↑
Avg. AUC ↑
Δ AUC
Syco ↓
Length ↓
Code ↓
BCE only
0.77
0.74
0.78
0.76
−0.06
40.6
181
26.2
Contrastive only
0.73
0.71
0.75
0.73
−0.09
44.2
204
34.8
BCE + contrastive
0.82
0.79
0.84
0.82
0.00
36.4
157
18.4
No projection head
0.72
0.70
0.73
0.72
−0.10
45.5
220
37.9
Frozen encoder only
0.70
0.68
0.71
0.70
−0.12
47.8
238
42.6
Appendix
Table 9: Auditor objective ablations. The full BCE+contrastive objective is substantially stronger than either loss alone. Contrastive learning improves both detection AUC and downstream mitigation, especially on code gaming.
Hyperparameter
Value
Avg. AUC ↑
Syco ↓
Length ↓
Code ↓
Default
–
0.82
36.4
157
18.4
τC
0.03
0.74
43.6
207
33.1
τC
0.05
0.78
40.8
185
27.4
τC
0.10
0.82
36.4
157
18.4
τC
0.20
0.81
37.2
155
20.1
τC
0.30
0.79
39.0
170
23.5
Appendix
Table 10: Contrastive-loss hyperparameters. Moderate contrastive weighting is important. Removing the contrastive term or making it too strong both degrade Auditor quality and downstream mitigation.
Hacker objective
Exploit yield ↑
Final hack ↓
Log-evasion: Rθ+λlogAξ
0.71
36.4
Gated reward: RθAξγ
0.44
39.8
Additive evasion: Rθ−λ(1−Aξ)
0.63
37.6
Appendix
Table 11: Hacker-objective ablation. Log-evasion produces the highest verified exploit yield and the lowest final hacking rate after AG-RLHF.
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at https://github.com/THUAIS-Lab/CATCH.
Shouli Wang, Yanfeng Jia, Zhihao Ou +6
Tsinghua University · Beihang University · Southern University of Science and Technology +2
Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.
Xuekang Wang, Zhuoyuan Hao, Shuo Hou +3
Tsinghua University · Harbin Institute of Technology, Shenzhen · Xi’an Jiaotong University
Reward models are central to large language model (LLM) alignment, but they remain vulnerable to reward hacking. To evaluate reward-model robustness, we introduce RewardHackBench containing 13 reward-hacking patterns covering real life high-stakes domains and general settings, and we find severe failures on specific subcategories across eight reward models. To mitigate these failures, we propose HARVE, a training-free reward-head editing method for scalar reward models. Instead of fine-tuning the reward model, HARVE identifies a multi-directional hacking subspace from residual stream directions associated with selected hacking subcategories, and removes the component of the reward-head vector aligned with that subspace. This directly reduces the reward head's sensitivity to hacking-related features using only a small set of contrastive gold-hacked examples, without gradient updates or fine-tuning. Comprehensive experiments across eight reward models indicates that \model improves hacking robustness, outperforms fine-tuning baselines, and preserves reward-models' general capability. Further analyses suggest that reward hacking is better captured as a multidimensional residual-space structure than by isolated surface cues.
Shuang Liu, Yuxuan Bo, Qiuyang Zhao +4
Carnegie Mellon University · University of Virginia · Harvard University +4