cs.LGFeb 4, 2026

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

Authors: Xiao Zhu, Xinyu Zhou, Boyu Zhu, Hanxu Hu, Mingzhe Du, Haotian Zhang, Huiming Wang, Zhijiang Guo

Organizations: LARK, HKUST(GZ) · Kuaishou Technology · UCL · UZH · NUS · HKUST

Abstract

Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose CodeScaler, a reward model designed to scale both reinforcement learning training and test-time inference for code generation. CodeScaler is trained on carefully curated preference data derived from verified code problems and incorporates syntax-aware code extraction and validity-preserving reward shaping to ensure stable and robust optimization. Across four coding benchmarks, CodeScaler consistently outperforms execution-based RL by +1.55 points on Qwen3-8B-Base and +4.23 points on Qwen3-14B-Base. By further scaling to 44K problems with additional synthetic data, CodeScaler yields +14.64 points improvement over the base model without requiring any test cases. At inference time, CodeScaler serves as an effective test-time scaling method, achieving performance comparable to unit test approaches while providing a 10-fold reduction in latency. Moreover, CodeScaler surpasses existing reward models on RM-Bench not only in the code domain (+3.3 points), but also in general and reasoning domains (+2.7 points on average).

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study

    Jun 18, 2026Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2Code GenerationReinforcement Learning With Verifiable Reward

  2. ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models

    Apr 30, 2026Jiasheng Zheng, Xin Zheng, Boxi Cao +8Large Language Model TrainingLarge-Scale Benchmark