Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose CodeScaler, a reward model designed to scale both reinforcement learning training and test-time inference for code generation. CodeScaler is trained on carefully curated preference data derived from verified code problems and incorporates syntax-aware code extraction and validity-preserving reward shaping to ensure stable and robust optimization. Across four coding benchmarks, CodeScaler consistently outperforms execution-based RL by +1.55 points on Qwen3-8B-Base and +4.23 points on Qwen3-14B-Base. By further scaling to 44K problems with additional synthetic data, CodeScaler yields +14.64 points improvement over the base model without requiring any test cases. At inference time, CodeScaler serves as an effective test-time scaling method, achieving performance comparable to unit test approaches while providing a 10-fold reduction in latency. Moreover, CodeScaler surpasses existing reward models on RM-Bench not only in the code domain (+3.3 points), but also in general and reasoning domains (+2.7 points on average).
Figures & tables
Figure 1 : Left: Training-Time Comparison shows that despite their larger data scale, synthetic code datasets exhibit a clear performance gap compared to verified code contest problems in RL training. While reward models provide dense supervision, they do not integrate effectively with RL, resulting in weaker performance than RLVR. Right: Test-Time Comparison illustrates that Unit Test TTS methods and off-the-shelf reward models demonstrate a clear performance–latency trade-off. This motivates us to develop a reward model that is both effective and efficient for RL training and test-time scaling.
Figure 2 : Overall Pipeline of CodeScaler for training-time and test-time scaling, which provides execution-free rewards for policy optimization during RL training, and serves as a lightweight sampler for Best-of-N selection, without relying on test-case execution.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
DeepCoder
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ RLVR
23.60
32.00
76.01
16.43
37.01
w/ CodeScaler (Ours)
24.80 (1.20 ↑ )
33.94 (1.94 ↑ )
75.90 (0.11 ↓ )
19.61 (3.27 ↑ )
38.56 (1.55 ↑ )
Qwen3-14B-Base
21.37
26.25
71.09
8.45
31.79
Table 1: Evaluation results ( Avg@8 ) across four code benchmarks, comparing CodeScaler with RLVR on DeepCoder, KodCode, and rStarCoder datasets. Bold indicates the best results. ↑(↓) indicates the improvement (degradation) compared to the RLVR baseline.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
DeepCoder
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ RLVR
23.60
32.00
76.01
16.43
37.01
w/ SkyworkRM
18.50
23.22
67.59
8.00
29.33 (7.68 ↓ )
w/ AceCodeRM
22.75
28.34
71.94
9.74
33.19 (3.82 ↓ )
Table 2: Comparison of CodeScaler with other reward models on the DeepCoder dataset ( Avg@8 ). Bold indicates the best results, Underline denotes the second-best results. ↑(↓) indicates the improvement (degradation) compared to the RLVR baseline.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ SkyworkRM †
17.69
23.48
75.22
7.52
30.98
w/ AceCodeRM
18.81
27.56
74.15
9.20
32.43
w/ CodeScaler (Ours)
25.76 (12.01 ↑ )
33.63 (14.60 ↑ )
76.81 (15.11 ↑ )
22.18 (16.83 ↑ )
39.60 (14.64 ↑ )
† Best checkpoint before training collapse.
Table 3: Evaluation results ( Avg@8 ) with scaled training on DeepCoder + synthetic data (44K problems). ↑(↓) indicates the improvement (degradation) compared to the base model.
Figure 3 : Comparison of Best-of-N ( BoN@8 ) performance across four code generation benchmarks using different test-time scaling methods. CodeScaler consistently outperforms other reward models and achieves performance comparable to CURE.
Time(s)
CURE
CodeScaler
Unit Test Gen.
979.3
N/A
Execution
516.7
N/A
RM Compute
N/A
146.1
Avg(s/question).
3.20
0.31
Table 4: Latency comparison between CURE and CodeScaler .
Models
RM-Bench
Avg.
Chat
Math
Code
Safety
Hard
Normal
Easy
SkyworkRM-8B
80.6
75.0
73.6
96.5
67.0
85.5
91.8
81.4
AceCodeRM-7B
66.7
65.3
66.9
89.9
62.2
74.4
79.9
72.2
AceCodeRM-32B
73.7
70.5
72.1
88.0
65.5
78.3
78.3
76.1
CodeScaler-8B (Ours)
83.0 (2.4 ↑ )
79.9 (4.9 ↑ )
76.9 (3.3 ↑ )
96.4 (0.1 ↓ )
71.8 (4.8 ↑ )
87.9 (2.4 ↑ )
92.5 (0.7 ↑ )
84.1 (2.7 ↑ )
Table 5: Evaluation results on RM-Bench across diverse reward models. Bold indicates the best results. ↑(↓) indicates the improvement (degradation) compared to the SkyworkRM-8B.
Language
Pass@1
SkyworkRM BoN@8
CodeScaler BoN@8
Python
30.7
34.1
34.8 (4.1 ↑ )
C++
28.0
32.9
33.2 (5.2 ↑ )
Java
26.8
32.9
34.8 (8.0 ↑ )
JavaScript
24.4
25.6
25.6 (1.2 ↑ )
Go
21.4
24.4
25.0 (3.6 ↑ )
Table 6 : Multi-lingual BoN@8 evaluation on HumanEval-X using Qwen3-8B as the policy model. ↑(↓) indicates the improvement (degradation) compared to the Pass@1.
Traj.
Benchmarks (Avg@8)
LiveCodeBench
CodeContests
CodeForces
rStarCoder
KodCode
23.79
28.19
11.58
DeepCoder
24.50 (0.71 ↑ )
30.75 (2.56 ↑ )
15.95 (4.37 ↑ )
Table 7 : Ablation on CodeScaler training data. Results compare RL training performance using different trajectory sources.
Figure 4 : Ablation study on CodeScaler components. Results compare RL performance using different reward model variants.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
DeepCoder
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ RLVR (Binary)
23.60
32.00
76.01
16.43
37.01
w/ CodeScaler (Binary)
24.80 (1.20 ↑ )
33.94 (1.94 ↑ )
75.90 (0.11 ↓ )
19.61 (3.27 ↑ )
38.56 (1.55 ↑ )
w/ CodeScaler (Pass Ratio)
24.20 (0.60 ↑ )
35.04 (3.04 ↑ )
71.09 (4.92 ↓ )
21.44 (5.01 ↑ )
37.94 (0.93 ↑ )
Table 8 : Ablation study on preference pair construction. Results compare RL training performance using different CodeScaler variant. Bold indicates the best results, Underline denotes the second-best results. ↑(↓) indicates the improvement (degradation) compared to the RLVR (Binary) baseline.
RM Variants
Benchmarks (BoN@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
CodeScaler (Binary)
41.29
36.82
80.09
13.33
42.88
CodeScaler (Pass Ratio)
40.84
35.88
79.64
13.28
42.41
Table 9: Ablation study on preference pair construction. Results compare Best-of-N sampling performance using different CodeScaler variant. Bold indicates the best results.
RM Variants
Benchmarks
BoN@8
RM-Bench
Avg.
Avg.
Avg.
SkyworkRM -8B
29.33
41.17
81.46
AceCodeRM -7B
33.19 (3.86 ↑ )
40.48 (0.69 ↓ )
72.20 (9.26 ↓ )
CodeScaler -1.7B
35.40 (6.07 ↑ )
38.16 (3.01 ↓ )
78.83 (2.63 ↓ )
CodeScaler -4B
37.20 (7.87 ↑ )
39.82 (1.35 ↓ )
82.86 (1.40 ↑ )
CodeScaler -8B
38.56 (9.23 ↑ )
44.15 (3.10 ↑ )
84.05 (2.59 ↑ )
Table 10: Ablation study on CodeScaler in diverse scales. Evaluation across code benchmarks, BoN sampling and RM-Bench. ↑(↓) indicates the improvement (degradation) compared to the SkyworkRM-8B.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Models
RM-Bench
Avg.
Chat
Math
Code
Safety
Hard
Normal
Easy
AceCodeRM-7B
66.7
65.3
66.9
89.9
62.2
74.4
79.9
72.2
AceCodeRM-32B
73.7
70.5
72.1
88.0
65.5
78.3
78.3
76.1
SkyworkRM-1.7B
69.6
71.4
72.3
92.9
54.5
82.3
92.8
76.6
CodeScaler-1.7B
74.4 (4.8 ↑ )
74.7 (3.3 ↑ )
73.1 (0.8 ↑ )
93.1 (0.2 ↑ )
61.5 (7.0 ↑ )
83.2 (0.9 ↑ )
91.7 (1.1 ↓ )
78.8 (2.2 ↑ )
SkyworkRM-4B
78.2
73.6
74.4
95.7
64.4
85.0
92.1
80.5
Appendix
Table 11: Full evaluation results on RM-Bench across diverse reward models. ↑(↓) indicates the improvement (degradation) compared to the SkyworkRM at same model size.
Policy
Reward Model
Pass@1
BoN@8
DeepSeek-Coder-6.7B-Instruct
SkyworkRM
18.41
21.76
AceCodeRM
23.85
CodeScaler
25.10 (6.69 ↑ )
Llama-3.1-8B-Instruct
SkyworkRM
15.90
19.67
AceCodeRM
20.08
CodeScaler
21.76 (5.86 ↑ )
Appendix
Table 12 : Best-of- N ( BoN@8 ) evaluation on CodeContests with non-Qwen policy models. ↑(↓) indicates the improvement (degradation) compared to the Pass@1.
Figure 5 : Training dynamics of invalid rate and fragment rate for Qwen3-8B-Base and Qwen3-14B-Base trained with CodeScaler on DeepCoder. Both rates decrease steadily after initial exploration, demonstrating effective suppression of failure modes.
Figure 6 : Evolution of the RM score distribution at training steps 0, 100 and 200 ( Qwen3-8B-Base on DeepCoder). The distribution shifts rightward steadily with no signs of reward hacking.
Figure 7 : Extended training of Qwen3-8B-Base on DeepCoder from 250 to 650 steps. Both RM score and pass@1 increase steadily with no divergence, confirming the absence of reward hacking.
Misalignment Ratio
Score on Misaligned Code
Margin (pos − neg)
AUROC
0%
−6.87±2.71
15.42±5.69
0.911
30%
−14.85±2.31
23.40±5.37
0.990
50%
−17.45±2.20
25.68±6.11
0.990
Appendix
Table 13 : Sensitivity analysis of the misalignment augmentation ratio. The 0% → 30% jump yields a large improvement in discrimination (AUROC 0.911 → 0.990), while the 30% → 50% gain is marginal, confirming that 30% is a robust and sufficient augmentation ratio.
Figure 8 : Both solutions implement an O(n) single-pass approach. The FN (correct, 278 chars) sweeps left-to-right, counting for each white ball the number of black balls to its left. The TN solution (incorrect, 661 chars) mirrors the traversal right-to-left, inadvertently solving the opposite goal (blacks left, whites right). Despite being 2.4× shorter and correct, the FN receives RM = 3.969 vs. RM = 7.531 for the wrong solution.
Reinforcement learning with verifiable rewards (RLVR) trains language models using programmatically checkable signals such as unit-test outcomes, enabling direct optimization for functional correctness in code generation. We conduct an empirical study of RLVR for Python code generation on the MBPP benchmark using two small models (Qwen3-0.6B and Llama3.2-1B) with LoRA fine-tuning. Across multiple reward formulations such as: unit-test-only rewards, static-analysis-only shaping via the Ruff linter, and a combined reward, we compare group-based policy optimization variants (GRPO and GSPO) and evaluate both functional correctness and behavioral diagnostics. In our experimental setting, RLVR improves pass@1 on MBPP test by up to 13 percentage points under proposed combined reward configuration. However, we find that reward shaping can induce systematic behavioral shifts: using only static-analysis penalties may bias the policy toward shorter completions that reduce lint errors without reliably improving functional correctness. In contrast, combined rewards mitigate this degeneration and yield more stable trade-offs between correctness and style constraints. Overall, our results highlight that RLVR effectiveness for code generation is highly sensitive to reward design and optimization granularity, and that diagnostics beyond pass@1, including generation length, Ruff severity profiles, and execution error types are useful for identifying failure modes.
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or infeasible, especially in coding tasks. Recent work addresses this by deriving rewards from a model's own signals, such as majority voting or confidence-based scores, achieving notable success on mathematical reasoning benchmarks. However, code generation poses distinct challenges: programs are structurally complex, semantically equivalent solutions may differ syntactically, and verification typically requires execution. Whether these intrinsic reward methods transfer effectively to code remains unexplored. In this work, we present a systematic empirical study of intrinsic reward methods for code generation. We conduct extensive experiments on LiveCodeBench, systematically evaluating representative certainty-based Reinforcement Learning from Internal Feedback (RLIF) approaches under different training scenarios and hyperparameter settings. Our experiments reveal that certainty-based methods yield early gains but inevitably collapse: models progressively shorten outputs and lose reasoning capability, with collapse speed sensitive to sample size and temperature. When used to initialize RLVR training, RLIF pre-training offers no significant improvement over training from scratch. We also provide actionable recommendations for using intrinsic rewards for training code reasoning models. Our study shows both the promise and limitations of intrinsic reward methods for code, informing future work on code models and agents.
Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2
Purdue University · UC Berkeley · UC Santa Barbara
Code sandboxes have emerged as a critical infrastructure for advancing the coding capabilities of large language models, providing verifiable feedback for both RL training and evaluation. However, existing systems fail to provide accurate verification and efficiency under high-concurrency workloads. We present ScaleBox, a high-fidelity and scalable system designed to address these limitations in large-scale code training. ScaleBox introduces automated special-judge generation and management, fine-grained parallel execution across test cases with seamless multi-node coordination, and a configuration-driven evaluation suite for reproducible benchmarking. A series of experiments demonstrates that ScaleBox significantly enhances code verification accuracy and efficiency. Our further RLVR experiments show that ScaleBox substantially improves both performance on LiveCodeBench and training stability, significantly outperforming heuristic-matching baselines. By providing a reliable and high-throughput infrastructure, ScaleBox facilitates more effective research and development in large-scale code training.
Jiasheng Zheng, Xin Zheng, Boxi Cao +8
1Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences · 1Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences