Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
Figures & tables
Overlap
Reward
Source
Low
High
Low
High
Δrˉ (pp)
Qwen3-4B-Base
Correct
29.1
4.1
0.578
0.390
{\color[rgb]{0.0859,0.5273,0.4219}-18.8}
Incorrect
27.8
7.3
0.070
0.163
{\color[rgb]{0.7461,0.2344,0.1758}+9.3}
Qwen3-8B-Base
Correct
29.6
4.3
0.625
0.418
{\color[rgb]{0.0859,0.5273,0.4219}-20.7}
Table 1 : Continuation resampling by original response outcome and window entropy. Δrˉ denotes the reward difference (High − Low).
Method
AIME24
AIME25
AIME26
HMMT26
AMC23
MATH500-H
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Qwen3-4B-Base
10.7
43.3
5.6
30.0
6.8
23.3
3.9
24.2
42.3
90.0
43.0
82.8
GRPO
12.5
33.3
10.4
36.7
7.5
33.3
5.9
24.2
54.5
90.0
56.1
85.1
EntropyAdv
13.1
46.7
12.5
36.7
10.0
33.3
6.3
27.3
53.8
90.0
56.7
86.6
HAPO
12.2
36.7
11.0
40.0
8.9
30.0
6.5
24.2
52.0
87.5
53.2
85.1
80/20
11.9
33.3
12.7
40.0
8.6
30.0
6.6
24.2
54.1
95.0
52.8
84.3
Table 2 : Results of diverse RLVR methods with base and reasoning models across six mathematical reasoning benchmarks. Bold denotes the best performance within each backbone.
Method
Graph
Cube
Sudoku
Family
K&K
Anag.
Zebra
Palin.
Avg.
Qwen3-4B-Base
6.75
24.63
40.00
17.00
29.13
37.50
29.50
9.00
24.19
GRPO
7.50
28.13
41.88
25.38
38.63
43.50
41.88
11.38
29.78
EntropyAdv
8.38
27.00
37.00
27.00
41.00
45.38
41.88
14.50
30.27
HAPO
10.13
27.13
40.38
30.13
43.63
47.88
47.50
14.13
32.61
80/20
11.38
26.75
39.25
31.50
49.63
46.38
47.38
15.13
33.43
RLRT
7.38
26.75
38.50
26.50
43.38
42.88
44.13
13.63
30.39
Table 3 : OOD generalization on Reasoning Gym tasks (avg@16). Bold denotes best performance.
Method
Hnorm ( ↑ )
Collision ( ↓ )
GRPO
0.657
0.166
EntropyAdv
0.659
0.159
HAPO
0.647
0.173
80/20
0.672
0.161
RLRT
0.659
0.168
EAPO (Ours)
0.751
0.101
Table 4 : Normalized answer entropy and collision rate on low-accuracy problems.
Method
Hnorm ( ↑ )
Collision ( ↓ )
GRPO
0.657
0.166
EntropyAdv
0.659
0.159
HAPO
0.647
0.173
80/20
0.672
0.161
RLRT
0.659
0.168
EAPO (Ours)
0.751
0.101
Table 4 : Normalized answer entropy and collision rate on low-accuracy problems.
Entropy
Penalization ( A^i<0 ): b−
Preference
Low ( −1 )
Uniform ( 0 )
High ( +1 )
Reinforcement ( A^i>0 ): b+
High ( +1 )
31.00 / 61.27
26.73 / 53.83
25.36 / 52.09
Uniform ( 0 )
26.08 / 51.26
24.48 / 50.43
22.89 / 50.03
Low ( −1 )
23.99 / 49.84
23.07 / 48.62
22.56 / 47.83
Table 5 : Ablation on sign–entropy coupling. Macro average of avg@32 / pass@32 (%) across six benchmarks. Our default EAPO configuration is shaded . Best scores are bolded .
κ
AIME25
AIME26
0
10.42 / 36.67
7.50 / 33.33
log2
10.73 / 33.33
10.63 / 26.67
log4
17.08 / 46.67
15.73 / 50.00
log8
18.23 / 46.67
15.83 / 40.00
log16
15.73 / 43.33
16.15 / 46.67
Table 6 : Sensitivity to κ . We report avg@32 / pass@32 (%) across different values of κ .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
High entropy
Low entropy
Qwen3-4B-Base
47.29±6.20
53.73±7.11
Qwen3-8B-Base
51.44±8.02
49.86±6.42
Appendix
Table 7 : Starting positions of selected windows (% of response length; mean ± std.).
Category
Parameter
Value
Data
Maximum response length
10,240 (base); 38,912 (reasoning)
Soft overlong penalty onset
8,192 (base); 32,768 (reasoning)
Batching
Prompts per rollout iteration
16
Responses per prompt ( G )
8
Generation batch size
128
Mini-batch size
64
Appendix
Table 8 : Training hyperparameters shared by EAPO and the baselines. Parameters not listed here follow the defaults of the original implementations.
Entropy
Penalization ( A^i<0 ): b−
Preference
Low ( −1 )
Uniform ( 0 )
High ( +1 )
Reinforcement ( A^i>0 ): b+
High ( +1 )
34.02 / 65.55
32.14 / 63.62
29.77 / 52.40
Uniform ( 0 )
30.55 / 59.52
26.78 / 52.18
27.48 / 52.56
Low ( −1 )
27.80 / 54.68
25.26 / 50.27
24.63 / 49.82
Appendix
Table 9 : Sign–entropy coupling on Qwen3-8B-Base. Macro average of avg@32 / pass@32 (%) across six benchmarks. Our default EAPO configuration is shaded . Best scores are bolded .
Method
Mode
AIME24
AIME25
AIME26
HMMT26
AMC23
MATH500-H
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
GRPO
LoRA
12.5
33.3
10.4
36.7
7.5
33.3
5.9
24.2
54.5
90.0
56.1
85.1
Full
13.1
36.7
12.8
36.7
10.2
33.3
6.9
27.3
53.7
92.5
58.2
87.3
EntropyAdv
LoRA
13.1
46.7
12.5
36.7
10.0
33.3
6.3
27.3
53.8
90.0
56.7
86.6
Full
13.0
36.7
11.0
36.7
10.0
36.7
7.7
30.3
52.5
87.5
57.2
84.3
HAPO
LoRA
12.2
36.7
11.0
40.0
8.9
30.0
6.5
24.2
52.0
87.5
53.2
85.1
Appendix
Table 10 : LoRA vs. full fine-tuning on Qwen3-4B-Base across mathematical reasoning benchmarks.
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR methods typically rely on on-policy optimization from scratch, resulting in high sampling costs and inefficient utilization of accumulated experience. As model capabilities and policy behaviors evolve during training, recent attempts to reuse experience via fixed reasoning trajectories further suffer from policy mismatch. Motivated by these limitations, we argue that experience in RLVR should not be reused as fixed reasoning trajectories, but instead expressed in a policy-adaptive manner. In this work, we propose Experience-Augmented Policy Optimization (EAPO), which leverages a prior RL-optimized policy as an action-level experience prior and selectively injects experience at critical decision points during rollout. To ensure stable and unbiased learning from experience-augmented rollouts, EAPO further incorporates an adapted importance sampling scheme. Experiments on using Qwen-2.5-math 7b and Qwen-3-8B on five different benchmarks demonstrate that EAPO consistently improves reasoning performance over state-of-the-art RLVR methods.
Jinda Lu, Kexin Huang, Junkang Wu +7
University of Science and Technology of China · 2Peking University · 3Dartmouth College +1
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult. Existing scalable RL methods typically assign trajectory-level rewards uniformly across tokens, while recent entropy-aware approaches either rely on coarse detached heuristics or directly optimize true entropy, which can introduce non-local gradient components misaligned with sampled-token policy updates. We propose Adaptive Credit Policy Optimization (ACPO), a token-level credit assignment framework based on a mode-local surrogate entropy. ACPO asymmetrically modulates policy updates by emphasizing uncertain decisions in successful rollouts and overconfident tokens in failed rollouts. We show that the surrogate admits deterministic entropy bounds and, under modal alignment and proximal updates, preserves the policy-gradient direction to leading order. Experiments on mathematical reasoning and coding benchmarks, including AIME 2025 and HumanEvalPro, show that ACPO consistently improves over strong RL baselines such as DAPO, GTPO, and SAPO.
Reinforcement learning with verifiable rewards has become the standard recipe for improving LLM reasoning, but the dominant algorithm GRPO assigns a single trajectory-level advantage to every token, diluting the signal at pivotal reasoning steps and injecting noise at uninformative ones. Critic-free alternatives derived from on-policy distillation supply per-token signals through oracle-conditioned likelihood ratios, yet apply each signal in isolation from the trajectory-level evidence accumulated up to that position. We propose Oracle-Prompted Policy Optimization (OPPO), which rests on a single observation: the oracle signal used by prior distillation-style methods for local discrimination is also the natural Bayesian update of the model's belief about eventual success. Accumulating the signal along a trajectory yields, in closed form and at the cost of one extra forward pass, a running estimate of the success probability at every position, together with a token-level advantage that requires no learned value network and no additional rollouts. A first-order analysis factorizes the advantage into the per-token discrimination signal used by distillation methods modulated by a state weight that concentrates credit on genuinely pivotal tokens, with a directional variance-reduction guarantee. The framework admits two estimators differing only in which model scores the evidence: a \textit{self-oracle} that reuses the student and recovers the on-policy distillation reward as a strict special case, and a \textit{teacher-oracle} that delegates scoring to a stronger frozen model. On two base LLMs across seven mathematics, science, and code reasoning benchmarks, OPPO improves over GRPO, DAPO, and SDPO by up to +6.0 points on AMC'23 and +5.2 points on AIME'24, with gains that widen monotonically with response length.
Yu Li, Rui Miao, Tian Lan +1
1George Washington University · 2The University of Texas at Dallas