cs.AISep 30, 2026

Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization

Authors: Yun Kim, Nojun Kwak

Organizations: Seoul National University

Abstract

Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR

    Jun 7, 2026Shumeng Yang, Yisu Liu, Jiayi Zheng +2Reinforcement Learning With Verifiable RewardToken-Level Entropy

  2. Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index

    Jun 30, 2026Outongyi Lv, Yanzhao Zheng, Yuanwei Zhang +5Token-Level EntropyReinforcement Learning With Verifiable Reward

  3. EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance

    May 6, 2026Song Yu, Li Li, Wenwen Zhao +1Reinforcement Learning With Verifiable RewardFlow-Grpo