Organizations: School of Computer Science, Peking University · National Key Laboratory for Multimedia Information Processing, Peking University · Foundation Model Department, Tencent
Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose πPPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, πPPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that πPPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.
Figures & tables
Method
Mathematical Reasoning
Tokens
AIME 24 Avg@16
AIME 25 Avg@16
Beyond AIME Avg@8
HMMT Feb Avg@16
HMMT Nov Avg@16
Overall Avg.
Gen./ Train 109
(a) Qwen3-4B
Initial policy
21.7
18.9
10.8
12.3
7.3
14.2
–
Critic-free
GRPO
60.4
54.8
34.5
34.6
41.5
45.2
2.33 / 2.38
DAPO
63.1
56.7
34.5
34.6
41.9
46.2
4.42 / 2.64
Table 1: Mathematical reasoning accuracy (%). “ Avg@k ” denotes mean accuracy over k random generations (i.e., pass@1); Overall is the mean over the five benchmarks. Best result within each model size is in bold.
Method
Overall
Qwen3-4B Avg.
Qwen3-8B Avg.
π PPO
50.3
51.6
w/o labels
44.1
48.8
w/ same polarity
43.6
48.6
w/ GT only
45.6
48.5
Table 2: Effects of correctness labels, reference polarity, and ground-truth-only context on average accuracy (%) across five benchmarks.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Qwen3-4B
Qwen3-8B
GPQA Avg@8
MMLU Avg@1
Overall Avg.
GPQA Avg@8
MMLU Avg@1
Overall Avg.
Initial policy
45.7
60.8
53.3
51.0
66.1
58.6
Critic-free
GRPO
52.9
67.3
60.1
58.3
71.6
65.0
DAPO
52.8
67.7
60.3
56.0
72.0
64.0
Actor–critic
Appendix
Table 4: Out-of-domain accuracy (%) of policies trained on DAPO-17K. “ Avg@k ” denotes mean accuracy over k random generations (i.e., pass@1); Overall is the mean over the two benchmarks. Asym: the method above with a smaller critic (0.6B for 4B, 1.7B for 8B). Best result within each model size is in bold.
Method
Mathematical Reasoning
AIME 24 Avg@16
AIME 25 Avg@16
Beyond AIME Avg@8
HMMT Feb Avg@16
HMMT Nov Avg@16
Overall Avg.
(a) Qwen3-4B
Initial policy
21.7
18.9
10.8
12.3
7.3
14.2
RLSD
57.1
47.9
31.5
31.5
38.8
41.4
π PPO (ours)
66.7
61.5
39.0
37.5
46.9
50.3
(b) Qwen3-8B
Appendix
Table 5: Comparison with RLSD on mathematical reasoning benchmarks. Overall averages the five benchmark scores; bold denotes the best result within each model size.