Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
Figures & tables
Overlap
Reward
Source
Low
High
Low
High
Δrˉ (pp)
Qwen3-4B-Base
Correct
29.1
4.1
0.578
0.390
{\color[rgb]{0.0859,0.5273,0.4219}-18.8}
Incorrect
27.8
7.3
0.070
0.163
{\color[rgb]{0.7461,0.2344,0.1758}+9.3}
Qwen3-8B-Base
Correct
29.6
4.3
0.625
0.418
{\color[rgb]{0.0859,0.5273,0.4219}-20.7}
Table 1 : Continuation resampling by original response outcome and window entropy. Δrˉ denotes the reward difference (High − Low).
Method
AIME24
AIME25
AIME26
HMMT26
AMC23
MATH500-H
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Qwen3-4B-Base
10.7
43.3
5.6
30.0
6.8
23.3
3.9
24.2
42.3
90.0
43.0
82.8
GRPO
12.5
33.3
10.4
36.7
7.5
33.3
5.9
24.2
54.5
90.0
56.1
85.1
EntropyAdv
13.1
46.7
12.5
36.7
10.0
33.3
6.3
27.3
53.8
90.0
56.7
86.6
HAPO
12.2
36.7
11.0
40.0
8.9
30.0
6.5
24.2
52.0
87.5
53.2
85.1
80/20
11.9
33.3
12.7
40.0
8.6
30.0
6.6
24.2
54.1
95.0
52.8
84.3
Table 2 : Results of diverse RLVR methods with base and reasoning models across six mathematical reasoning benchmarks. Bold denotes the best performance within each backbone.
Method
Graph
Cube
Sudoku
Family
K&K
Anag.
Zebra
Palin.
Avg.
Qwen3-4B-Base
6.75
24.63
40.00
17.00
29.13
37.50
29.50
9.00
24.19
GRPO
7.50
28.13
41.88
25.38
38.63
43.50
41.88
11.38
29.78
EntropyAdv
8.38
27.00
37.00
27.00
41.00
45.38
41.88
14.50
30.27
HAPO
10.13
27.13
40.38
30.13
43.63
47.88
47.50
14.13
32.61
80/20
11.38
26.75
39.25
31.50
49.63
46.38
47.38
15.13
33.43
RLRT
7.38
26.75
38.50
26.50
43.38
42.88
44.13
13.63
30.39
Table 3 : OOD generalization on Reasoning Gym tasks (avg@16). Bold denotes best performance.
Method
Hnorm ( ↑ )
Collision ( ↓ )
GRPO
0.657
0.166
EntropyAdv
0.659
0.159
HAPO
0.647
0.173
80/20
0.672
0.161
RLRT
0.659
0.168
EAPO (Ours)
0.751
0.101
Table 4 : Normalized answer entropy and collision rate on low-accuracy problems.
Method
Hnorm ( ↑ )
Collision ( ↓ )
GRPO
0.657
0.166
EntropyAdv
0.659
0.159
HAPO
0.647
0.173
80/20
0.672
0.161
RLRT
0.659
0.168
EAPO (Ours)
0.751
0.101
Table 4 : Normalized answer entropy and collision rate on low-accuracy problems.
Entropy
Penalization ( A^i<0 ): b−
Preference
Low ( −1 )
Uniform ( 0 )
High ( +1 )
Reinforcement ( A^i>0 ): b+
High ( +1 )
31.00 / 61.27
26.73 / 53.83
25.36 / 52.09
Uniform ( 0 )
26.08 / 51.26
24.48 / 50.43
22.89 / 50.03
Low ( −1 )
23.99 / 49.84
23.07 / 48.62
22.56 / 47.83
Table 5 : Ablation on sign–entropy coupling. Macro average of avg@32 / pass@32 (%) across six benchmarks. Our default EAPO configuration is shaded . Best scores are bolded .
κ
AIME25
AIME26
0
10.42 / 36.67
7.50 / 33.33
log2
10.73 / 33.33
10.63 / 26.67
log4
17.08 / 46.67
15.73 / 50.00
log8
18.23 / 46.67
15.83 / 40.00
log16
15.73 / 43.33
16.15 / 46.67
Table 6 : Sensitivity to κ . We report avg@32 / pass@32 (%) across different values of κ .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
High entropy
Low entropy
Qwen3-4B-Base
47.29±6.20
53.73±7.11
Qwen3-8B-Base
51.44±8.02
49.86±6.42
Appendix
Table 7 : Starting positions of selected windows (% of response length; mean ± std.).
Category
Parameter
Value
Data
Maximum response length
10,240 (base); 38,912 (reasoning)
Soft overlong penalty onset
8,192 (base); 32,768 (reasoning)
Batching
Prompts per rollout iteration
16
Responses per prompt ( G )
8
Generation batch size
128
Mini-batch size
64
Appendix
Table 8 : Training hyperparameters shared by EAPO and the baselines. Parameters not listed here follow the defaults of the original implementations.
Entropy
Penalization ( A^i<0 ): b−
Preference
Low ( −1 )
Uniform ( 0 )
High ( +1 )
Reinforcement ( A^i>0 ): b+
High ( +1 )
34.02 / 65.55
32.14 / 63.62
29.77 / 52.40
Uniform ( 0 )
30.55 / 59.52
26.78 / 52.18
27.48 / 52.56
Low ( −1 )
27.80 / 54.68
25.26 / 50.27
24.63 / 49.82
Appendix
Table 9 : Sign–entropy coupling on Qwen3-8B-Base. Macro average of avg@32 / pass@32 (%) across six benchmarks. Our default EAPO configuration is shaded . Best scores are bolded .
Method
Mode
AIME24
AIME25
AIME26
HMMT26
AMC23
MATH500-H
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
Avg@32
Pass@32
GRPO
LoRA
12.5
33.3
10.4
36.7
7.5
33.3
5.9
24.2
54.5
90.0
56.1
85.1
Full
13.1
36.7
12.8
36.7
10.2
33.3
6.9
27.3
53.7
92.5
58.2
87.3
EntropyAdv
LoRA
13.1
46.7
12.5
36.7
10.0
33.3
6.3
27.3
53.8
90.0
56.7
86.6
Full
13.0
36.7
11.0
36.7
10.0
36.7
7.7
30.3
52.5
87.5
57.2
84.3
HAPO
LoRA
12.2
36.7
11.0
40.0
8.9
30.0
6.5
24.2
52.0
87.5
53.2
85.1
Appendix
Table 10 : LoRA vs. full fine-tuning on Qwen3-4B-Base across mathematical reasoning benchmarks.