cs.AIApr 1, 2026

RefineRL: Advancing Competitive Programming with Self-Refinement Reinforcement Learning

Authors: Shaopeng Fu, Xingxing Zhang, Li Dong, Furu Wei

Organizations: Work done during an internship at Microsoft Research · KAUST · Microsoft Research

Abstract

While large language models (LLMs) have demonstrated strong performance on complex reasoning tasks such as competitive programming (CP), existing methods predominantly focus on single-attempt settings, overlooking their capacity for iterative refinement. In this paper, we present RefineRL, a novel approach designed to unleash the self-refinement capabilities of LLMs for CP problem solving. RefineRL introduces two key innovations: (1) Skeptical-Agent, an iterative self-refinement agent equipped with local execution tools to validate generated solutions against public test cases of CP problems. This agent always remains skeptical of its own outputs and thereby enforces rigorous self-refinement even when validation suggests correctness. (2) A reinforcement learning (RL) solution to incentivize LLMs to self-refine with only standard RLVR data (i.e., problems paired with their verifiable answers). Extensive experiments on Qwen3-4B and Qwen3-4B-2507 demonstrate that our method yields substantial gains: after our RL training, these compact 4B models integrated with the Skeptical-Agent not only outperform much larger 32B models but also approach the single-attempt performance of 235B models. Notably, although our models are trained solely on CP coding data, RefineRL also substantially improves their performance on mathematical reasoning benchmarks, suggesting that the learned self-refinement behavior can transfer beyond coding-specific tasks. These findings highlight self-refinement as a promising direction for scaling LLMs' general reasoning ability.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback

    May 18, 2026Anika Tabassum, Md Sifat Hossain, Md. Fahim Arefin +2Code GenerationFeedback

  2. X-Coder: Advancing Competitive Programming with Synthetic Tasks, Solutions, and Tests

    Jan 11, 2026Jie Wu, Haoling Li, Xin Zhang +8Raw Judge OutputsSynthetic Task

  3. Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution

    May 14, 2026Han Li, Jinyu Tian, Rili Feng +10Self-Evolving AgentsTransferable Implicit Solvent Machine Learning Potential