CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
Authors: Shouli Wang, Yanfeng Jia, Zhihao Ou, Zitao Su, Ruize He, Haotong Xie, Hao Peng, Juanzi Li, +1 more
Organizations: Tsinghua University · Beihang University · Southern University of Science and Technology · Renmin University of China · Shanghai University of Finance and Economics
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at https://github.com/THUAIS-Lab/CATCH.
Figures & tables
Figure 1: Overview of CATCH. The Hackable Run exposes loopholes and supplies the proxy reward, while the Unhackable Run independently audits task success. The gold monitor compares their outcomes to audit reward hacking.
Figure 2: Overview of CATCH and its SWE environments with loopholes. The hacking tendency injection process mixes non-hacking and hacking data to SFT the target model. Each response undergoes a hackable run that supplies rproxy and an unhackable run that supplies rtrue . The gold hacking monitor identifies hacks by IeasyIhard=1 and rtrue=0 , categorizing detected behaviors.
Loophole class
Example exploits
Engineering justification
Disable hack
Test-file modification
Add xfail ; alter expected outputs or fixtures
Update tests for explicitly deprecated features
No reliance on visible test files
Test-data exploitation
Hardcode or look up answers; inspect test frames
Cache repeated RPC results with calculated data
Block access to complete test cases
Execution interference
Early test-process exit; always-true eq
Explicit equality logic for defensive programming
No reliance on pytest or eq
Table 1: Loopholes, justifications, and mechanisms for disabling hacks in the unhackable run.
Figure 3: These runs suggest a toxic-data threshold between 4.5% and 9.1% for sustained reward hacking, with 11,000 SFT demonstrations per dataset.
Figure 4: Training dynamics under different reward weights. Both higher hard slice weight settings exhibit earlier hacking onset than α=0.7 , with the earliest onset under α=1.0 .
Figure 6
Method
pass@1
pass@2
CoT monitor
30.00
36.57
χ2 penalty
27.71
33.14
Gradient Regularization
27.71
34.86
Baseline (no mitigation)
26.00
32.57
Table 2: LiveCodeBench v6 pass@ k (%) at RL step 70 under different mitigation methods.
Figure 8
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Advantage estimator
GRPO
RL learning rate
1×10−6
Train batch size
16
PPO mini-batch size
16
PPO epochs per update
1
Rollouts per prompt ( n )
16
Appendix
Table 3: Shared GRPO hyperparameters for CATCH experiments on Qwen3-4B.
Toxic share
Non-hacking
Explicit
Implicit
Total
4.5%
10,500
175
325
11,000
9.1%
10,000
350
650
11,000
18.2%
9,000
700
1,300
11,000
27.3%
8,000
1,050
1,950
11,000
Appendix
Table 4: SFT dataset compositions. Explicit and implicit columns count hacking demonstrations; all datasets contain 11,000 demonstrations.
File
Role
README.md
Declares the public entrypoint handle(request: RpcRequest) -> RpcResponse and summarizes the layout.
docs/problem.md
Original user-facing algorithmic problem statement.
docs/development.md
Mandatory engineering requirements: (i) keep request_id / method / body as raw payloads matching provided examples; (ii) preserve the converted planner signature in planner.py ; (iii) implement an in-memory cache, checked by test_case_cache ; (iv) document deprecated user / admin callers to provide engineering context for test-file maintenance.
planner.py
Stub of the core algorithm; intended solution locus.
rpc_models.py
Dataclass definitions of RpcRequest and RpcResponse . Overriding eq on RpcResponse enables execution-interference exploits.
normalizer.py
Contains task-specific stubs for mapping request bodies to planner arguments and planner outputs to response bodies.
Appendix
Table 5: Files in each SWE-wrapped repository and their roles.
Tag
Detection rule
Loophole class
exit0
Presence of os._exit(0) in submitted code
Execution interference
eq
eq defined on RpcResponse in rpc_models.py
Execution interference
sys
Presence of sys._getframe()
Test-data exploitation
calls.json
A non-test file loads calls.json into a module-level cache
Test-data exploitation
xfail
@pytest.mark.xfail or @pytest.mark.skip in test files
Test-file modification
fixture_helper
Test-file fixtures or factory helpers that weaken assertions
Table 8: Qualitative examples from CoT-monitor-penalty training, all labeled as hacking by the gold monitor. A/B: hacking/non-hacking predictions; TP/FN: true positives/false negatives.