cs.LGSep 29, 2026

Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

Authors: Kun Liang, Chenming Tang, Clive Bai, Weijie Liu, Zeyuan Liu, Qingyang Zhang, Saiyong Yang, Yunfang Wu

Organizations: School of Computer Science, Peking University · National Key Laboratory for Multimedia Information Processing, Peking University · Foundation Model Department, Tencent

Abstract

Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose ππPPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, ππPPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that ππPPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

    May 8, 2026Yunho Choi, Jongwon Lim, Woojin Ahn +3Reinforcement Learning With Verifiable RewardVerifiable Rewards

  2. Start Classifying: Categorical Critics for LLM Reinforcement Learning

    Aug 3, 2026Zhijian Zhou, Long Li, Xuan Zhang +7Large Language Model Reinforcement LearningCritic-Free Reinforcement Learning

  3. VIMPO: Value-Implicit Policy Optimization for LLMs

    Jun 18, 2026Zhewei Kang, Aosong Feng, Sergey Levine +2Reinforcement Learning With Verifiable RewardFrictive Policy Optimization