cs.AIFeb 2, 2026

Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking

Authors: Mohammad Beigi, Ming Jin, Junshan Zhang, Qifan Wang, Lifu Huang

Organizations: Department of Computer Science, University of California, Davis, USA · Virginia Tech, Blacksburg, USA · Meta AI Research.

Abstract

Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores while violating human intent. Existing mitigations rely on static defenses that cannot adapt to novel exploitation strategies. We propose Adversarial Reward Auditing (ARA), a framework that reconceptualizes reward hacking as a dynamic, competitive game. ARA operates in two stages: first, a Hacker policy discovers reward model vulnerabilities while an Auditor learns to detect exploitation from latent representations; second, Auditor-Guided RLHF (AG-RLHF) gates reward signals to penalize detected hacking, transforming reward hacking from an unobservable failure into a measurable, controllable signal. Experiments across three hacking scenarios demonstrate that ARA achieves the best alignment-utility tradeoff among all baselines: reducing sycophancy to near-SFT levels while improving helpfulness, decreasing verbosity while achieving the highest ROUGE-L, and suppressing code gaming while improving Pass@1. Beyond single-domain evaluation, we show that reward hacking, detection, and mitigation all generalize across domains -- a Hacker trained on code gaming exhibits increased sycophancy despite no reward for this behavior, and an Auditor trained on one domain effectively suppresses exploitation in others, enabling efficient multi-domain defense with a single model.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL

    Sep 30, 2026Shouli Wang, Yanfeng Jia, Zhihao Ou +6Reinforcement Learning With Verifiable RewardOffline Reinforcement Learning

  2. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

    Jun 3, 2026Xuekang Wang, Zhuoyuan Hao, Shuo Hou +3Reward Signal

  3. HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models

    Jun 2, 2026Shuang Liu, Yuxuan Bo, Qiuyang Zhao +4Progress Reward ModelingLarge Language Model Alignment