CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
Authors: Shouli Wang, Yanfeng Jia, Zhihao Ou, Zitao Su, Ruize He, Haotong Xie, Hao Peng, Juanzi Li, +1 more
Organizations: Tsinghua University · Beihang University · Southern University of Science and Technology · Renmin University of China · Shanghai University of Finance and Economics
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at https://github.com/THUAIS-Lab/CATCH.
Figures & tables
Figure 1: Overview of CATCH. The Hackable Run exposes loopholes and supplies the proxy reward, while the Unhackable Run independently audits task success. The gold monitor compares their outcomes to audit reward hacking.
Figure 2: Overview of CATCH and its SWE environments with loopholes. The hacking tendency injection process mixes non-hacking and hacking data to SFT the target model. Each response undergoes a hackable run that supplies rproxy and an unhackable run that supplies rtrue . The gold hacking monitor identifies hacks by IeasyIhard=1 and rtrue=0 , categorizing detected behaviors.
Loophole class
Example exploits
Engineering justification
Disable hack
Test-file modification
Add xfail ; alter expected outputs or fixtures
Update tests for explicitly deprecated features
No reliance on visible test files
Test-data exploitation
Hardcode or look up answers; inspect test frames
Cache repeated RPC results with calculated data
Block access to complete test cases
Execution interference
Early test-process exit; always-true eq
Explicit equality logic for defensive programming
No reliance on pytest or eq
Table 1: Loopholes, justifications, and mechanisms for disabling hacks in the unhackable run.
Figure 3: These runs suggest a toxic-data threshold between 4.5% and 9.1% for sustained reward hacking, with 11,000 SFT demonstrations per dataset.
Figure 4: Training dynamics under different reward weights. Both higher hard slice weight settings exhibit earlier hacking onset than α=0.7 , with the earliest onset under α=1.0 .
Figure 6
Method
pass@1
pass@2
CoT monitor
30.00
36.57
χ2 penalty
27.71
33.14
Gradient Regularization
27.71
34.86
Baseline (no mitigation)
26.00
32.57
Table 2: LiveCodeBench v6 pass@ k (%) at RL step 70 under different mitigation methods.
Figure 8
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Advantage estimator
GRPO
RL learning rate
1×10−6
Train batch size
16
PPO mini-batch size
16
PPO epochs per update
1
Rollouts per prompt ( n )
16
Appendix
Table 3: Shared GRPO hyperparameters for CATCH experiments on Qwen3-4B.
Toxic share
Non-hacking
Explicit
Implicit
Total
4.5%
10,500
175
325
11,000
9.1%
10,000
350
650
11,000
18.2%
9,000
700
1,300
11,000
27.3%
8,000
1,050
1,950
11,000
Appendix
Table 4: SFT dataset compositions. Explicit and implicit columns count hacking demonstrations; all datasets contain 11,000 demonstrations.
File
Role
README.md
Declares the public entrypoint handle(request: RpcRequest) -> RpcResponse and summarizes the layout.
docs/problem.md
Original user-facing algorithmic problem statement.
docs/development.md
Mandatory engineering requirements: (i) keep request_id / method / body as raw payloads matching provided examples; (ii) preserve the converted planner signature in planner.py ; (iii) implement an in-memory cache, checked by test_case_cache ; (iv) document deprecated user / admin callers to provide engineering context for test-file maintenance.
planner.py
Stub of the core algorithm; intended solution locus.
rpc_models.py
Dataclass definitions of RpcRequest and RpcResponse . Overriding eq on RpcResponse enables execution-interference exploits.
normalizer.py
Contains task-specific stubs for mapping request bodies to planner arguments and planner outputs to response bodies.
Appendix
Table 5: Files in each SWE-wrapped repository and their roles.
Tag
Detection rule
Loophole class
exit0
Presence of os._exit(0) in submitted code
Execution interference
eq
eq defined on RpcResponse in rpc_models.py
Execution interference
sys
Presence of sys._getframe()
Test-data exploitation
calls.json
A non-test file loads calls.json into a module-level cache
Test-data exploitation
xfail
@pytest.mark.xfail or @pytest.mark.skip in test files
Test-file modification
fixture_helper
Test-file fixtures or factory helpers that weaken assertions
Table 8: Qualitative examples from CoT-monitor-penalty training, all labeled as hacking by the gold monitor. A/B: hacking/non-hacking predictions; TP/FN: true positives/false negatives.
Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task. Precisely measuring reward hacking occurrence remains challenging because true task rewards are often expensive or impossible to compute. We introduce Countdown-Code, a minimal environment where models can both solve a mathematical reasoning task and manipulate the test harness. This dual-access design creates a clean separation between proxy rewards (test pass/fail) and true rewards (mathematical correctness), enabling accurate measurement of reward-hacking rates. Using this environment, we study reward hacking in open-weight LLMs and find that such behaviors can be unintentionally learned during supervised fine-tuning (SFT) when even a small fraction of reward-hacking trajectories leak into training data. As little as 1% contamination in distillation SFT data is sufficient for models to internalize reward hacking which resurfaces during subsequent reinforcement learning (RL). We further show that RL amplifies misalignment and drives its generalization beyond the original domain. We open-source our environment and code to facilitate future research on reward hacking in LLMs. Our results reveal a previously underexplored pathway through which reward hacking can emerge and persist in LLMs, underscoring the need for more rigorous validation of synthetic SFT data. Code is available at https://github.com/zohaib-khan5040/Countdown-Code.
Muhammad Khalifa, Zohaib Khan, Omer Tafveez +2
NVIDIA · University of Michigan · University of Illinois Urbana-Champaign
Reward hacking in code generation, where models exploit evaluation loopholes to obtain high reward without correctly solving the intended task, poses a critical challenge for Reinforcement Learning (RL) and the deployment of reasoning models. Existing studies often rely on explicitly prompted hacking trajectories, but it remains unclear whether monitors trained on such data can detect reward hacks that arise without direct hacking instructions during RL training. In this work, we introduce Trace-and-Amplify, a framework for scalable curation of reward-hacking trajectories that arise during RL training without explicit hacking instructions. The framework uses unit-test tracers to identify hacking solutions when they occur and retains such trajectories for monitor training and evaluation. Through controlled comparisons between monitors trained on prompt-elicited hacking trajectories and training-time reward-hacking trajectories collected by Trace-and-Amplify, we find that \textbf{(1) prompt-elicited-data-trained monitors often fail to generalize to trajectories curated by our framework}, and \textbf{(2) monitors trained on our Trace-and-Amplify trajectories demonstrate stronger generalizability to unseen hacking types}. Our results indicate that prompted reward hacking data may not fully reflect training-time reward-hacking behaviors, and that relying solely on these data can lead to misleading conclusions. Codebase is available at https://github.com/LichenLillc/CoTMonitoring.git
Lichen Li, Hengguang Zhou, Yijun Liang +2
Peking University · University of California, Los Angeles · University of Maryland +2
Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.
Xuekang Wang, Zhuoyuan Hao, Shuo Hou +3
Tsinghua University · Harbin Institute of Technology, Shenzhen · Xi’an Jiaotong University