cs.AIOct 1, 2026

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Authors: Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao

Organizations: National University of Singapore · The Hong Kong University of Science and Technology · University of Science and Technology of China

Abstract

Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SCOPE-RL: Optimizing Reasoning Paths Before and After Success

    Jul 13, 2026Xiaojian Liu, Han Xu, Jianqiang Xia +6Reinforcement Learning With Verifiable RewardVerifiable Rewards

  2. Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

    Sep 29, 2026Hongyang Li, Xiao Li, Caesar Wu +3Critic-Free Reinforcement LearningReward-Only Policy Optimization Variants