Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores while violating human intent. Existing mitigations rely on static defenses that cannot adapt to novel exploitation strategies. We propose Adversarial Reward Auditing (ARA), a framework that reconceptualizes reward hacking as a dynamic, competitive game. ARA operates in two stages: first, a Hacker policy discovers reward model vulnerabilities while an Auditor learns to detect exploitation from latent representations; second, Auditor-Guided RLHF (AG-RLHF) gates reward signals to penalize detected hacking, transforming reward hacking from an unobservable failure into a measurable, controllable signal. Experiments across three hacking scenarios demonstrate that ARA achieves the best alignment-utility tradeoff among all baselines: reducing sycophancy to near-SFT levels while improving helpfulness, decreasing verbosity while achieving the highest ROUGE-L, and suppressing code gaming while improving Pass@1. Beyond single-domain evaluation, we show that reward hacking, detection, and mitigation all generalize across domains -- a Hacker trained on code gaming exhibits increased sycophancy despite no reward for this behavior, and an Auditor trained on one domain effectively suppresses exploitation in others, enabling efficient multi-domain defense with a single model.
Figures & tables
Figure 1: ARA framework. A Hacker exploits a frozen reward model while an Auditor gates RL reward.
Scenario
Model
Metric
SFT
PPO
PPO+KL
ODIN
InFoRM
RM Ens.
Adv-RM
RM-FT †
ARA
Sycophancy
Llama-2-7B
Rate ↓
36.2
71.8
57.6
52.3
48.1
51.9
51.4
46.8
39.1
Help ↑
41.3
77.4
67.8
64.1
62.6
65.0
70.6
68.3
76.8
Win ↑
–
60.7
61.4
62.8
61.9
62.1
64.2
63.1
69.4
Llama-3.1-8B
Rate ↓
34.1
67.6
53.8
48.4
43.7
49.8
47.6
43.1
36.4
Help ↑
44.8
79.2
70.1
67.2
65.2
66.9
72.8
70.4
78.6
Win ↑
–
62.3
62.8
64.1
63.7
64.1
66.2
64.8
71.8
Table 1: Main results across three reward-hacking scenarios and three model families. For each scenario we report one alignment metric ( ↓ ), one utility metric ( ↑ ), and GPT-4o win rate against SFT ( ↑ ). Bold indicates the best result among RLHF methods; underlining indicates the runner-up.
Figure 2: Proxy reward versus true task performance during RLHF on Llama-3.1-8B.
Figure 3: Proxy reward Rθ(x,y) vs. alignment with the preference direction ux for Llama-3.1-8B. Genuine responses and verified exploits achieve comparable proxy rewards (horizontal overlap) but separate along ux (vertical axis).
Variant
ux ?
AUC ↑
Syco. ↓
Len. ↓
Code ↓
Random neg.
No
0.58
49.2
238
39.6
Direct ux probe
Yes
0.64
48.6
229
37.4
Shuffled ux in S
Corrupt
0.66
44.1
214
34.0
C -only
Yes
0.70
41.0
181
25.8
No filter
No
0.74
40.3
186
27.9
G -only
No
0.78
38.1
169
21.7
Table 2: Hard-negative construction ablation. AUC is averaged across the three scenarios.
Scenario
Direct ux
C -only
G -only
Full ARA
Syco.
0.41
0.55
0.74
0.79
Len.
0.43
0.57
0.72
0.77
Code
0.46
0.60
0.76
0.81
Avg.
0.43
0.57
0.74
0.79
Table 3: ux -adversarial detection. AUC when the usual association between ux -alignment and exploitation is reversed.
Detection Method
Gen.
Expl.
Δ
Proxy Reward ( Rθ )
0.91
0.89
0.02
GPT-4
0.78
0.34
0.44
Auditor Confidence
0.89
0.21
0.68
Table 4 : Genuine vs. exploit detection. Mean scores for 500 genuine and 500 exploit responses ( Rθ>0.85 ).
Hacker Attack (%)
Auditor AUC
ARA Defense
Train
Syco.
Len.
Code
Syco.
Len.
Code
Syco. ↓
Len. ↓
Code ↓
Ref.
36.2 s
8.1 s
4.2 s
–
–
–
72.4 p
347 p
61.3 p
Syco.
72.4
35.2
28.3
0.82
0.77
0.58
38.4
178
48.2
Length
41.8
52.3
22.6
0.75
0.79
0.55
42.1
162
51.6
Code
58.7
31.4
61.3
0.62
0.59
0.84
54.8
246
19.6
All
68.9
78.1
59.7
–
–
–
39.2
165
21.4
Table 5: Cross-domain attack, detection, and defense. Left: Hacker attack rate ( s SFT baseline). Middle: Auditor AUC. Right: ARA defense metrics ( p PPO baseline).
Figure 4: Gating severity sensitivity analysis. Hacking metrics (red, lower is better) and utility metrics (blue, higher is better).
Figure 5: Gating impact on genuine vs. exploitative responses (Llama-3.1-8B, Rθ>0.85 ). Top: Auditor score distributions. Bottom: proxy vs. gated reward.
Error (%)
Retention (%)
Rank ρ
FPR
FNR
Genuine
Exploit
Genuine
Exploit
Syco. ( γ =2)
6.8
8.6
81.8
6.1
.58
− .04
Len. ( γ =2)
4.4
6.8
78.8
5.4
.72
− .15
Code ( γ =3)
5.6
5.2
80.5
7.3
.75
− .25
Table 6: Gating impact at γ∗ (Llama-3.1-8B). FPR/FNR: misclassification rates. Retention: % of proxy reward preserved after gating. ρ : Spearman correlation between Rθ and Rgated .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Hyperparameter
Value
Hacker (PPO)
Learning rate
1×10−6
Batch size
128
PPO epochs
4
Clip ratio ϵ
0.2
Evasion weight λA
0.5
KL penalty βH
0.1
Appendix
Table 7: Hyperparameters.
Figure 6: Auditor size ablation across all scenarios. Top row: Detection AUC. Bottom row: Mitigation performance (red = hacking metric ↓ , blue = utility metric ↑ ). Gold markers indicate optimal configuration for each task. Different tasks require different capacities: length bias saturates at 25M, sycophancy at 35M, and code gaming at 50M, reflecting increasing exploitation complexity.
Hyperparameter
Value
Syco ↓
Help ↑
Length ↓
ROUGE-L ↑
Code ↓
Pass@1 ↑
Avg. AUC ↑
Default ARA
–
36.4
78.6
157
25.0
18.4
38.4
0.82
λA
0.1
43.0
76.5
190
24.2
29.6
36.8
0.75
λA
0.3
38.9
78.4
165
24.8
21.7
38.1
0.80
λA
0.5
36.4
78.6
157
25.0
18.4
38.4
0.82
λA
0.7
35.8
76.9
154
24.7
17.9
37.2
0.83
λA
1.0
34.9
73.9
151
23.9
17.5
35.8
0.82
Appendix
Table 8: Stage-1 hyperparameter sensitivity. The default is not the best value for every individual metric, but it provides the strongest overall alignment–utility tradeoff. More aggressive Hacker pressure can lower hacking slightly, while usually reducing utility.
Auditor objective
Syco AUC ↑
Length AUC ↑
Code AUC ↑
Avg. AUC ↑
Δ AUC
Syco ↓
Length ↓
Code ↓
BCE only
0.77
0.74
0.78
0.76
−0.06
40.6
181
26.2
Contrastive only
0.73
0.71
0.75
0.73
−0.09
44.2
204
34.8
BCE + contrastive
0.82
0.79
0.84
0.82
0.00
36.4
157
18.4
No projection head
0.72
0.70
0.73
0.72
−0.10
45.5
220
37.9
Frozen encoder only
0.70
0.68
0.71
0.70
−0.12
47.8
238
42.6
Appendix
Table 9: Auditor objective ablations. The full BCE+contrastive objective is substantially stronger than either loss alone. Contrastive learning improves both detection AUC and downstream mitigation, especially on code gaming.
Hyperparameter
Value
Avg. AUC ↑
Syco ↓
Length ↓
Code ↓
Default
–
0.82
36.4
157
18.4
τC
0.03
0.74
43.6
207
33.1
τC
0.05
0.78
40.8
185
27.4
τC
0.10
0.82
36.4
157
18.4
τC
0.20
0.81
37.2
155
20.1
τC
0.30
0.79
39.0
170
23.5
Appendix
Table 10: Contrastive-loss hyperparameters. Moderate contrastive weighting is important. Removing the contrastive term or making it too strong both degrade Auditor quality and downstream mitigation.
Hacker objective
Exploit yield ↑
Final hack ↓
Log-evasion: Rθ+λlogAξ
0.71
36.4
Gated reward: RθAξγ
0.44
39.8
Additive evasion: Rθ−λ(1−Aξ)
0.63
37.6
Appendix
Table 11: Hacker-objective ablation. Log-evasion produces the highest verified exploit yield and the lowest final hacking rate after AG-RLHF.