Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose CodeScaler, a reward model designed to scale both reinforcement learning training and test-time inference for code generation. CodeScaler is trained on carefully curated preference data derived from verified code problems and incorporates syntax-aware code extraction and validity-preserving reward shaping to ensure stable and robust optimization. Across four coding benchmarks, CodeScaler consistently outperforms execution-based RL by +1.55 points on Qwen3-8B-Base and +4.23 points on Qwen3-14B-Base. By further scaling to 44K problems with additional synthetic data, CodeScaler yields +14.64 points improvement over the base model without requiring any test cases. At inference time, CodeScaler serves as an effective test-time scaling method, achieving performance comparable to unit test approaches while providing a 10-fold reduction in latency. Moreover, CodeScaler surpasses existing reward models on RM-Bench not only in the code domain (+3.3 points), but also in general and reasoning domains (+2.7 points on average).
Figures & tables
Figure 1 : Left: Training-Time Comparison shows that despite their larger data scale, synthetic code datasets exhibit a clear performance gap compared to verified code contest problems in RL training. While reward models provide dense supervision, they do not integrate effectively with RL, resulting in weaker performance than RLVR. Right: Test-Time Comparison illustrates that Unit Test TTS methods and off-the-shelf reward models demonstrate a clear performance–latency trade-off. This motivates us to develop a reward model that is both effective and efficient for RL training and test-time scaling.
Figure 2 : Overall Pipeline of CodeScaler for training-time and test-time scaling, which provides execution-free rewards for policy optimization during RL training, and serves as a lightweight sampler for Best-of-N selection, without relying on test-case execution.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
DeepCoder
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ RLVR
23.60
32.00
76.01
16.43
37.01
w/ CodeScaler (Ours)
24.80 (1.20 ↑ )
33.94 (1.94 ↑ )
75.90 (0.11 ↓ )
19.61 (3.27 ↑ )
38.56 (1.55 ↑ )
Qwen3-14B-Base
21.37
26.25
71.09
8.45
31.79
Table 1: Evaluation results ( Avg@8 ) across four code benchmarks, comparing CodeScaler with RLVR on DeepCoder, KodCode, and rStarCoder datasets. Bold indicates the best results. ↑(↓) indicates the improvement (degradation) compared to the RLVR baseline.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
DeepCoder
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ RLVR
23.60
32.00
76.01
16.43
37.01
w/ SkyworkRM
18.50
23.22
67.59
8.00
29.33 (7.68 ↓ )
w/ AceCodeRM
22.75
28.34
71.94
9.74
33.19 (3.82 ↓ )
Table 2: Comparison of CodeScaler with other reward models on the DeepCoder dataset ( Avg@8 ). Bold indicates the best results, Underline denotes the second-best results. ↑(↓) indicates the improvement (degradation) compared to the RLVR baseline.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ SkyworkRM †
17.69
23.48
75.22
7.52
30.98
w/ AceCodeRM
18.81
27.56
74.15
9.20
32.43
w/ CodeScaler (Ours)
25.76 (12.01 ↑ )
33.63 (14.60 ↑ )
76.81 (15.11 ↑ )
22.18 (16.83 ↑ )
39.60 (14.64 ↑ )
† Best checkpoint before training collapse.
Table 3: Evaluation results ( Avg@8 ) with scaled training on DeepCoder + synthetic data (44K problems). ↑(↓) indicates the improvement (degradation) compared to the base model.
Figure 3 : Comparison of Best-of-N ( BoN@8 ) performance across four code generation benchmarks using different test-time scaling methods. CodeScaler consistently outperforms other reward models and achieves performance comparable to CURE.
Time(s)
CURE
CodeScaler
Unit Test Gen.
979.3
N/A
Execution
516.7
N/A
RM Compute
N/A
146.1
Avg(s/question).
3.20
0.31
Table 4: Latency comparison between CURE and CodeScaler .
Models
RM-Bench
Avg.
Chat
Math
Code
Safety
Hard
Normal
Easy
SkyworkRM-8B
80.6
75.0
73.6
96.5
67.0
85.5
91.8
81.4
AceCodeRM-7B
66.7
65.3
66.9
89.9
62.2
74.4
79.9
72.2
AceCodeRM-32B
73.7
70.5
72.1
88.0
65.5
78.3
78.3
76.1
CodeScaler-8B (Ours)
83.0 (2.4 ↑ )
79.9 (4.9 ↑ )
76.9 (3.3 ↑ )
96.4 (0.1 ↓ )
71.8 (4.8 ↑ )
87.9 (2.4 ↑ )
92.5 (0.7 ↑ )
84.1 (2.7 ↑ )
Table 5: Evaluation results on RM-Bench across diverse reward models. Bold indicates the best results. ↑(↓) indicates the improvement (degradation) compared to the SkyworkRM-8B.
Language
Pass@1
SkyworkRM BoN@8
CodeScaler BoN@8
Python
30.7
34.1
34.8 (4.1 ↑ )
C++
28.0
32.9
33.2 (5.2 ↑ )
Java
26.8
32.9
34.8 (8.0 ↑ )
JavaScript
24.4
25.6
25.6 (1.2 ↑ )
Go
21.4
24.4
25.0 (3.6 ↑ )
Table 6 : Multi-lingual BoN@8 evaluation on HumanEval-X using Qwen3-8B as the policy model. ↑(↓) indicates the improvement (degradation) compared to the Pass@1.
Traj.
Benchmarks (Avg@8)
LiveCodeBench
CodeContests
CodeForces
rStarCoder
KodCode
23.79
28.19
11.58
DeepCoder
24.50 (0.71 ↑ )
30.75 (2.56 ↑ )
15.95 (4.37 ↑ )
Table 7 : Ablation on CodeScaler training data. Results compare RL training performance using different trajectory sources.
Figure 4 : Ablation study on CodeScaler components. Results compare RL performance using different reward model variants.
Models
Benchmarks (Avg@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
DeepCoder
Qwen3-8B-Base
13.75
19.03
61.70
5.35
24.96
w/ RLVR (Binary)
23.60
32.00
76.01
16.43
37.01
w/ CodeScaler (Binary)
24.80 (1.20 ↑ )
33.94 (1.94 ↑ )
75.90 (0.11 ↓ )
19.61 (3.27 ↑ )
38.56 (1.55 ↑ )
w/ CodeScaler (Pass Ratio)
24.20 (0.60 ↑ )
35.04 (3.04 ↑ )
71.09 (4.92 ↓ )
21.44 (5.01 ↑ )
37.94 (0.93 ↑ )
Table 8 : Ablation study on preference pair construction. Results compare RL training performance using different CodeScaler variant. Bold indicates the best results, Underline denotes the second-best results. ↑(↓) indicates the improvement (degradation) compared to the RLVR (Binary) baseline.
RM Variants
Benchmarks (BoN@8)
Avg.
LiveCodeBench
CodeContests
MBPP
CodeForces
CodeScaler (Binary)
41.29
36.82
80.09
13.33
42.88
CodeScaler (Pass Ratio)
40.84
35.88
79.64
13.28
42.41
Table 9: Ablation study on preference pair construction. Results compare Best-of-N sampling performance using different CodeScaler variant. Bold indicates the best results.
RM Variants
Benchmarks
BoN@8
RM-Bench
Avg.
Avg.
Avg.
SkyworkRM -8B
29.33
41.17
81.46
AceCodeRM -7B
33.19 (3.86 ↑ )
40.48 (0.69 ↓ )
72.20 (9.26 ↓ )
CodeScaler -1.7B
35.40 (6.07 ↑ )
38.16 (3.01 ↓ )
78.83 (2.63 ↓ )
CodeScaler -4B
37.20 (7.87 ↑ )
39.82 (1.35 ↓ )
82.86 (1.40 ↑ )
CodeScaler -8B
38.56 (9.23 ↑ )
44.15 (3.10 ↑ )
84.05 (2.59 ↑ )
Table 10: Ablation study on CodeScaler in diverse scales. Evaluation across code benchmarks, BoN sampling and RM-Bench. ↑(↓) indicates the improvement (degradation) compared to the SkyworkRM-8B.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Models
RM-Bench
Avg.
Chat
Math
Code
Safety
Hard
Normal
Easy
AceCodeRM-7B
66.7
65.3
66.9
89.9
62.2
74.4
79.9
72.2
AceCodeRM-32B
73.7
70.5
72.1
88.0
65.5
78.3
78.3
76.1
SkyworkRM-1.7B
69.6
71.4
72.3
92.9
54.5
82.3
92.8
76.6
CodeScaler-1.7B
74.4 (4.8 ↑ )
74.7 (3.3 ↑ )
73.1 (0.8 ↑ )
93.1 (0.2 ↑ )
61.5 (7.0 ↑ )
83.2 (0.9 ↑ )
91.7 (1.1 ↓ )
78.8 (2.2 ↑ )
SkyworkRM-4B
78.2
73.6
74.4
95.7
64.4
85.0
92.1
80.5
Appendix
Table 11: Full evaluation results on RM-Bench across diverse reward models. ↑(↓) indicates the improvement (degradation) compared to the SkyworkRM at same model size.
Policy
Reward Model
Pass@1
BoN@8
DeepSeek-Coder-6.7B-Instruct
SkyworkRM
18.41
21.76
AceCodeRM
23.85
CodeScaler
25.10 (6.69 ↑ )
Llama-3.1-8B-Instruct
SkyworkRM
15.90
19.67
AceCodeRM
20.08
CodeScaler
21.76 (5.86 ↑ )
Appendix
Table 12 : Best-of- N ( BoN@8 ) evaluation on CodeContests with non-Qwen policy models. ↑(↓) indicates the improvement (degradation) compared to the Pass@1.
Figure 5 : Training dynamics of invalid rate and fragment rate for Qwen3-8B-Base and Qwen3-14B-Base trained with CodeScaler on DeepCoder. Both rates decrease steadily after initial exploration, demonstrating effective suppression of failure modes.
Figure 6 : Evolution of the RM score distribution at training steps 0, 100 and 200 ( Qwen3-8B-Base on DeepCoder). The distribution shifts rightward steadily with no signs of reward hacking.
Figure 7 : Extended training of Qwen3-8B-Base on DeepCoder from 250 to 650 steps. Both RM score and pass@1 increase steadily with no divergence, confirming the absence of reward hacking.
Misalignment Ratio
Score on Misaligned Code
Margin (pos − neg)
AUROC
0%
−6.87±2.71
15.42±5.69
0.911
30%
−14.85±2.31
23.40±5.37
0.990
50%
−17.45±2.20
25.68±6.11
0.990
Appendix
Table 13 : Sensitivity analysis of the misalignment augmentation ratio. The 0% → 30% jump yields a large improvement in discrimination (AUROC 0.911 → 0.990), while the 30% → 50% gain is marginal, confirming that 30% is a robust and sufficient augmentation ratio.
Figure 8 : Both solutions implement an O(n) single-pass approach. The FN (correct, 278 chars) sweeps left-to-right, counting for each white ball the number of black balls to its left. The TN solution (incorrect, 661 chars) mirrors the traversal right-to-left, inadvertently solving the opposite goal (blacks left, whites right). Despite being 2.4× shorter and correct, the FN receives RM = 3.969 vs. RM = 7.531 for the wrong solution.
1Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences · 1Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences