Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
Organizations: Seoul National University
Abstract
Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.
Figures & tables
| Global Entropy | Proximal Entropy | |
| Hard | 0.315 | 0.00987 |
| Moderate | 0.270 | 0.00987 |
| Easy | 0.221 | 0.00989 |
| Model | Method | AIME2024 | AIME2025 | AMC | MATH500 | Mean |
| Qwen3-1.7B | Base | 14.02 0.82 | 17.58 0.74 | 45.21 0.65 | 75.53 0.48 | 38.09 0.42 |
| GRPO | 16.58 0.44 | 18.65 0.62 | 50.28 0.99 | 80.27 1.01 | 41.45 0.50 | |
| Entropy Adv. | 21.44 1.81 | 19.48 1.72 | 52.52 1.06 | 81.40 1.44 | 43.71 0.52 | |
| 80/20 | 20.97 1.48 | 20.83 0.75 | 52.13 0.44 | 80.73 1.50 | 43.67 0.30 | |
| PEPO (Ours) | 21.74 1.06 | 22.43 0.67 | 55.26 1.11 | 82.67 0.31 | 45.53 0.46 | |
| Qwen3-4B | Base | 25.62 0.95 | 20.14 0.81 | 55.39 0.72 | 79.80 0.54 | 45.24 0.51 |
| Model | Method | AIME2024 | AIME2025 | AMC | MATH500 | Mean |
| Qwen3-1.7B | 80/20 | 20.97 1.48 | 20.83 0.75 | 52.13 0.44 | 80.73 1.50 | 43.67 0.30 |
| 80/20 + Proximal Entropy | 22.01 0.87 | 20.69 0.24 | 53.36 0.78 | 82.27 0.31 | 44.58 0.43 | |
| PEPO (Ours) | 21.74 1.06 | 22.43 0.67 | 55.26 1.11 | 82.67 0.31 | 45.53 0.46 | |
| Qwen3-4B | 80/20 | 35.35 2.47 | 30.62 0.21 | 67.25 1.53 | 89.73 1.15 | 55.74 0.86 |
| 80/20 + Proximal Entropy | 36.18 2.61 | 31.60 0.48 | 69.38 1.41 | 90.33 0.31 | 56.87 0.56 | |
| PEPO (Ours) | 39.24 1.27 | 31.81 0.13 | 69.58 0.59 | 91.40 1.83 | 58.01 0.14 |
| Method | AIME2024 | AIME2025 | AMC | MATH500 | Mean |
| SPO | 35.70 1.97 | 31.39 0.44 | 70.81 1.28 | 90.60 0.53 | 57.13 0.42 |
| S-Entropy Adv. | 38.89 1.26 | 21.80 0.67 | 68.42 0.43 | 89.47 0.31 | 54.65 0.58 |
| S-80/20 | 35.28 1.15 | 27.64 1.05 | 70.06 0.93 | 91.00 0.40 | 55.99 0.43 |
| PESPO (Ours) | 41.74 0.52 | 31.60 0.79 | 71.71 1.06 | 90.67 0.31 | 58.93 0.48 |
| Budget | Global entropy | Proximal entropy |
| 0% | 25.62 | 25.62 |
| 5% | 27.34 | 31.62 |
| 10% | 29.91 | 33.33 |
| 15% | 30.76 | 33.33 |
| 20% | 32.47 | 33.33 |
| 30% | 33.33 | 33.33 |
| AIME2024 | AIME2025 | AMC | MATH500 | |
| 36.25 1.50 | 30.62 1.30 | 68.39 0.59 | 90.67 0.42 | |
| 39.24 1.27 | 31.81 0.13 | 69.58 0.59 | 91.40 1.83 | |
| 36.81 0.24 | 28.27 1.26 | 69.35 0.49 | 92.27 0.46 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
| Framework | ROLL |
| Training data | DeepMath-103K |
| Rollouts per prompt ( ) | 8 |
| Learning rate | |
| Optimizer | AdamW |
| Lower clipping threshold ( ) | 0.2 |
| Parameter | Value |
| Framework | VeRL |
| Training data | DAPO-Math-17K |
| Rollouts per prompt | 1 (single-stream) |
| Learning rate | |
| Optimizer | AdamW |
| Lower clipping threshold ( ) | 0.2 |
| Method | Qwen3-1.7B | Qwen3-4B |
| GRPO | 1d 7h 32m | 1d 19h 25m |
| 80/20 | 1d 7h 28m | 1d 20h 14m |
| Entropy Adv. | 1d 8h 28m | 1d 21h 3m |
| PEPO (Ours) | 1d 7h 54m | 1d 19h 45m |
| Method | Correct | Pass@1 | Gain vs. GRPO |
| GRPO | 335 / 1,055 | 31.75% | — |
| 80/20 | 433 / 1,055 | 41.04% | +9.29 pp |
| PEPO | 489 / 1,055 | 46.35% | +14.60 pp |
| Subset | Problems | GRPO | 80/20 | PEPO | PEPO GRPO |
| By difficulty | |||||
| Easy | 322 | 76.09% | 84.78% | 90.99% | +14.91 pp |
| Medium | 383 | 22.72% | 37.60% | 43.86% | +21.15 pp |
| Hard | 350 | 0.86% | 4.57% | 8.00% | +7.14 pp |
| By platform | |||||
| AtCoder | 602 | 31.23% | 38.87% | 42.86% | +11.63 pp |
| Configuration | Mean | Gain (pp) |
| Global entropy, binary selection (80/20) | 55.74 | — |
| Proximal entropy, binary selection | 56.87 | +1.14 |
| Proximal entropy, continuous weighting on all rollouts | 57.53 | +0.66 |
| Proximal entropy, continuous weighting on positive rollouts | 58.01 | +0.48 |
| AIME24 | AIME25 | AMC | MATH500 | Mean | |
| 0.5 | 8.75 | 1.25 | 23.72 | 49.40 | 20.78 |
| 1.0 | 9.10 | 1.46 | 24.12 | 49.33 | 21.00 |
| 2.0 | 8.54 | 1.46 | 23.72 | 49.00 | 20.68 |
| Model | Method | Run | AIME2024 | AIME2025 | AMC | MATH500 |
| Qwen3-1.7B | GRPO | 1 | 16.42 | 18.26 | 51.22 | 79.20 |
| 2 | 17.07 | 19.36 | 50.38 | 81.20 | ||
| 3 | 16.24 | 18.33 | 49.25 | 80.40 | ||
| Mean | 16.58 | 18.65 | 50.28 | 80.27 | ||
| Entropy Adv. | 1 | 22.25 | 20.53 | 51.61 | 79.80 | |
| 2 | 19.37 | 17.50 | 53.69 | 82.60 |
| Model | Method | Run | AIME2024 | AIME2025 | AMC | MATH500 |
| Qwen3-1.7B | 80/20 + Prox. | 1 | 22.71 | 20.83 | 53.16 | 82.60 |
| 2 | 21.04 | 20.42 | 52.71 | 82.20 | ||
| 3 | 22.29 | 20.83 | 54.22 | 82.00 | ||
| Mean | 22.01 | 20.69 | 53.36 | 82.27 | ||
| Qwen3-4B | 80/20 + Prox. | 1 | 38.75 | 31.04 | 69.65 | 90.60 |
| 2 | 33.54 | 31.88 | 70.63 | 90.40 |
| Method | Run | AIME2024 | AIME2025 | AMC | MATH500 |
| SPO | 1 | 37.92 | 31.25 | 70.71 | 90.00 |
| 2 | 35.00 | 31.04 | 72.14 | 90.80 | |
| 3 | 34.17 | 31.88 | 69.58 | 91.00 | |
| Mean | 35.70 | 31.39 | 70.81 | 90.60 | |
| S-80/20 | 1 | 36.46 | 26.67 | 68.98 | 90.60 |
| 2 | 35.21 | 28.75 | 70.56 | 91.40 |
| Window Size | Run | AIME2024 | AIME2025 | AMC | MATH500 |
| 1 | 34.59 | 30.20 | 68.36 | 90.20 | |
| 2 | 37.50 | 29.59 | 67.82 | 91.00 | |
| 3 | 36.67 | 32.08 | 68.99 | 90.80 | |
| Mean | 36.25 | 30.62 | 68.39 | 90.67 | |
| 1 | 37.08 | 27.08 | 69.87 | 92.00 | |
| 2 | 36.67 | 29.59 | 68.89 | 92.00 |