Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.
Figures & tables
Figure 1 : Validation accuracy over training steps for PEPO and the 80/20. The plotted values represent the average mean@k accuracy evaluated across the MATH, AIME 2024, AIME 2025, and AMC benchmarks for Qwen3-1.7B (a) and Qwen3-4B (b). PEPO achieves a higher validation accuracy compared to the 80/20 baseline.
Figure 2 : Positional bias in token entropy on Llama-3.2-3B-Instruct. We plot the deviation of global and proximal entropy from its sequence-level mean across generation percentiles.
Global Entropy
Proximal Entropy
Hard
0.315
0.00987
Moderate
0.270
0.00987
Easy
0.221
0.00989
Table 1: Prompt difficulty bias in token entropy on Llama-3.2-3B-Instruct.
Model
Method
AIME2024
AIME2025
AMC
MATH500
Mean
Qwen3-1.7B
Base
14.02 ± 0.82
17.58 ± 0.74
45.21 ± 0.65
75.53 ± 0.48
38.09 ± 0.42
GRPO
16.58 ± 0.44
18.65 ± 0.62
50.28 ± 0.99
80.27 ± 1.01
41.45 ± 0.50
Entropy Adv.
21.44 ± 1.81
19.48 ± 1.72
52.52 ± 1.06
81.40 ± 1.44
43.71 ± 0.52
80/20
20.97 ± 1.48
20.83 ± 0.75
52.13 ± 0.44
80.73 ± 1.50
43.67 ± 0.30
PEPO (Ours)
21.74 ± 1.06
22.43 ± 0.67
55.26 ± 1.11
82.67 ± 0.31
45.53 ± 0.46
Qwen3-4B
Base
25.62 ± 0.95
20.14 ± 0.81
55.39 ± 0.72
79.80 ± 0.54
45.24 ± 0.51
Table 2 : Main results on mathematical reasoning benchmarks. Trained methods are reported as mean ± sample standard deviation (SD) over three independent runs. We report avg@1 on MATH500 and avg@16 on AMC, AIME 2024, and AIME 2025; the Mean column summarizes each run’s four benchmark scores. Best results per model are in bold .
Model
Method
AIME2024
AIME2025
AMC
MATH500
Mean
Qwen3-1.7B
80/20
20.97 ± 1.48
20.83 ± 0.75
52.13 ± 0.44
80.73 ± 1.50
43.67 ± 0.30
80/20 + Proximal Entropy
22.01 ± 0.87
20.69 ± 0.24
53.36 ± 0.78
82.27 ± 0.31
44.58 ± 0.43
PEPO (Ours)
21.74 ± 1.06
22.43 ± 0.67
55.26 ± 1.11
82.67 ± 0.31
45.53 ± 0.46
Qwen3-4B
80/20
35.35 ± 2.47
30.62 ± 0.21
67.25 ± 1.53
89.73 ± 1.15
55.74 ± 0.86
80/20 + Proximal Entropy
36.18 ± 2.61
31.60 ± 0.48
69.38 ± 1.41
90.33 ± 0.31
56.87 ± 0.56
PEPO (Ours)
39.24 ± 1.27
31.81 ± 0.13
69.58 ± 0.59
91.40 ± 1.83
58.01 ± 0.14
Table 3 : Effect of substituting proximal entropy into 80/20. Each cell reports the mean ± sample SD over three independent runs. All other components of 80/20 remain unchanged. PEPO results from Table 2 are included for reference. Best results within each model are bolded.
Method
AIME2024
AIME2025
AMC
MATH500
Mean
SPO
35.70 ± 1.97
31.39 ± 0.44
70.81 ± 1.28
90.60 ± 0.53
57.13 ± 0.42
S-Entropy Adv.
38.89 ± 1.26
21.80 ± 0.67
68.42 ± 0.43
89.47 ± 0.31
54.65 ± 0.58
S-80/20
35.28 ± 1.15
27.64 ± 1.05
70.06 ± 0.93
91.00 ± 0.40
55.99 ± 0.43
PESPO (Ours)
41.74 ± 0.52
31.60 ± 0.79
71.71 ± 1.06
90.67 ± 0.31
58.93 ± 0.48
Table 4 : Results on single-stream RL (SPO) with Qwen3-4B. Cells show the mean followed by sample SD in smaller type, computed over three independent runs. All methods are implemented on top of SPO. PESPO uses the same hyperparameters as in the main experiments without modification.
Budget
Global entropy
Proximal entropy
0%
25.62
25.62
5%
27.34
31.62
10%
29.91
33.33
15%
30.76
33.33
20%
32.47
33.33
30%
33.33
33.33
Table 5 : Token-replacement intervention on AIME 2024 using Qwen3-4B. Results are avg@16 accuracy (%); each budget is the fraction of token positions replaced.
W
AIME2024
AIME2025
AMC
MATH500
W=51
36.25 ± 1.50
30.62 ± 1.30
68.39 ± 0.59
90.67 ± 0.42
W=101
39.24 ± 1.27
31.81 ± 0.13
69.58 ± 0.59
91.40 ± 1.83
W=151
36.81 ± 0.24
28.27 ± 1.26
69.35 ± 0.49
92.27 ± 0.46
Table 6: Ablation studies on window size on Qwen3-4B. Each cell shows the mean followed by sample SD in smaller type over three independent runs.
Figure 3 : Qualitative comparison of token weighting on a Qwen3-4B rollout. (a) Global entropy assigns broad, near-uniform weight across structural and decision tokens alike (e.g., the entire phrases “We’ll calculate” and “property of periodicity”). (b) Proximal entropy evaluates each token relative to its local neighborhood, producing fine-grained sub-phrase weighting that singles out the tokens that steer the reasoning trajectory (e.g., calculate , periodicity ) while still preserving signal across the derivation steps.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Framework
ROLL
Training data
DeepMath-103K
Rollouts per prompt ( G )
8
Learning rate
1×10−6
Optimizer
AdamW
Lower clipping threshold ( εl )
0.2
Appendix
Table 7 : Hyperparameters for main experiments on ROLL framework. These settings are used for GRPO, 80/20, Entropy Adv., and PEPO across all three models.
Parameter
Value
Framework
VeRL
Training data
DAPO-Math-17K
Rollouts per prompt
1 (single-stream)
Learning rate
1×10−6
Optimizer
AdamW
Lower clipping threshold ( εl )
0.2
Appendix
Table 8 : Hyperparameters for single-stream experiments on VeRL framework. These settings are used for SPO, S-80/20, S-Entropy Adv., and PESPO.
Method
Qwen3-1.7B
Qwen3-4B
GRPO
1d 7h 32m
1d 19h 25m
80/20
1d 7h 28m
1d 20h 14m
Entropy Adv.
1d 8h 28m
1d 21h 3m
PEPO (Ours)
1d 7h 54m
1d 19h 45m
Appendix
Table 9 : Total wall-clock training time on 2 × H200 GPUs. All methods are run with identical training configurations for 500 steps. PEPO adds negligible overhead over GRPO across both models.
Figure 4 : Deviation of windowed entropy statistics from their sequence-level mean across generation percentiles, for window sizes W∈{31,51,101} . We compare two aggregation choices: proximal entropy (solid purple), defined as a softmax over the window, and a simple windowed average (dashed orange). Both yield flat curves across all three window sizes, in contrast to global entropy (gray dashed) which decays substantially. This shows that the positional invariance is robust to the window size and is a property of the windowing itself rather than of the softmax operation. Computed over 992 rollouts on the validation set with Llama-3.2-3B-Instruct.
Method
Correct
Pass@1
Gain vs. GRPO
GRPO
335 / 1,055
31.75%
—
80/20
433 / 1,055
41.04%
+9.29 pp
PEPO
489 / 1,055
46.35%
+14.60 pp
Appendix
Table 10: Held-out LiveCodeBench release_v6/code_generation_lite pass@1 point estimates for Qwen3-4B.
Subset
Problems
GRPO
80/20
PEPO
PEPO − GRPO
By difficulty
Easy
322
76.09%
84.78%
90.99%
+14.91 pp
Medium
383
22.72%
37.60%
43.86%
+21.15 pp
Hard
350
0.86%
4.57%
8.00%
+7.14 pp
By platform
AtCoder
602
31.23%
38.87%
42.86%
+11.63 pp
Appendix
Table 11: LiveCodeBench pass@1 breakdown by difficulty and problem platform. The Codeforces subset contains only 9 problems.
Configuration
Mean
Gain (pp)
Global entropy, binary selection (80/20)
55.74
—
Proximal entropy, binary selection
56.87
+1.14
Proximal entropy, continuous weighting on all rollouts
57.53
+0.66
Proximal entropy, continuous weighting on positive rollouts
58.01
+0.48
Appendix
Table 12: Sequential weighting ablation on Qwen3-4B. Mean accuracy averages AIME 2024, AIME 2025, AMC, and MATH500; the first gain averages the benchmark-level differences in Table 3.
τ
AIME24
AIME25
AMC
MATH500
Mean
0.5
8.75
1.25
23.72
49.40
20.78
1.0
9.10
1.46
24.12
49.33
21.00
2.0
8.54
1.46
23.72
49.00
20.68
Appendix
Table 13: Temperature sensitivity for Llama-3.2-3B-Instruct. Results are the reported point estimates; no seed-level uncertainty was provided for this sweep. Best result per column is bolded.
Model
Method
Run
AIME2024
AIME2025
AMC
MATH500
Qwen3-1.7B
GRPO
1
16.42
18.26
51.22
79.20
2
17.07
19.36
50.38
81.20
3
16.24
18.33
49.25
80.40
Mean
16.58
18.65
50.28
80.27
Entropy Adv.
1
22.25
20.53
51.61
79.80
2
19.37
17.50
53.69
82.60
Appendix
Table 14 : Per-run accuracy for the main results in Table 2. Each row reports the avg@1 accuracy on MATH500 and avg@16 accuracy on AMC, AIME 2024, and AIME 2025 for a single run. Mean rows match the values reported in Table 2.
Model
Method
Run
AIME2024
AIME2025
AMC
MATH500
Qwen3-1.7B
80/20 + Prox.
1
22.71
20.83
53.16
82.60
2
21.04
20.42
52.71
82.20
3
22.29
20.83
54.22
82.00
Mean
22.01
20.69
53.36
82.27
Qwen3-4B
80/20 + Prox.
1
38.75
31.04
69.65
90.60
2
33.54
31.88
70.63
90.40
Appendix
Table 15 : Per-run accuracy for the drop-in proximal entropy results in Table 3. Mean rows match the values reported in Table 3.
Method
Run
AIME2024
AIME2025
AMC
MATH500
SPO
1
37.92
31.25
70.71
90.00
2
35.00
31.04
72.14
90.80
3
34.17
31.88
69.58
91.00
Mean
35.70
31.39
70.81
90.60
S-80/20
1
36.46
26.67
68.98
90.60
2
35.21
28.75
70.56
91.40
Appendix
Table 16 : Per-run accuracy for the single-stream RL results in Table 4 on Qwen3-4B. Mean rows match the values reported in Table 4.
Window Size
Run
AIME2024
AIME2025
AMC
MATH500
W=51
1
34.59
30.20
68.36
90.20
2
37.50
29.59
67.82
91.00
3
36.67
32.08
68.99
90.80
Mean
36.25
30.62
68.39
90.67
W=151
1
37.08
27.08
69.87
92.00
2
36.67
29.59
68.89
92.00
Appendix
Table 17 : Per-run accuracy for the window size ablation in Table 6 on Qwen3-4B. Mean rows match the values reported in Table 6.
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · School of Artificial Intelligence, Beijing University of Posts and Telecommunications +2