Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
Organizations: National University of Singapore · The Hong Kong University of Science and Technology · University of Science and Technology of China
Abstract
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
Figures & tables
| Method | Backbone | Verifier | MMLU-Pro | GPQA-Dia. | TheoremQA | WebInst. | MATH-500 | Minerva | AIME24 | All |
|---|---|---|---|---|---|---|---|---|---|---|
| Avg@2 | Avg@4 | Avg@2 | Avg@2 | Avg@2 | Avg@2 | Avg@16 | ||||
| Base | Qwen2.5-7B | - | 45.3 | 32.4 | 41.4 | 60.4 | 63.0 | 37.6 | 6.5 | 40.9 |
| RLVR | Rule | 55.1 | 36.2 | 52.2 | 75.3 | 76.5 | 54.9 | 17.7 | 52.6 | |
| Gen. Reasoner | Model | 55.4 | 37.4 | 52.1 | 74.5 | 77.0 | 51.7 | 16.0 | 52.0 | |
| VeriFree | Free | 53.8 | 36.7 | 47.6 | 72.5 | 73.5 | 49.0 | 12.5 | 49.4 | |
| RLPR | Free | 56.0 | 37.6 | 55.4 | 75.5 | 78.0 | 56.5 | 16.3 | 53.6 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.