TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
Organizations: University of Chinese Academy of Sciences · Meituan · MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences
Abstract
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
Figures & tables
| Method | AIME24 | AIME25 | AMC23 | MATH500 | Minerva | Olympiad | Average |
|---|---|---|---|---|---|---|---|
| Qwen3-14B | 50.4 | 40.0 | 79.8 | 87.9 | 46.1 | 60.3 | 60.8 |
| PPO | 55.2 | 45.8 | 90.6 | 93.9 | 49.1 | 63.5 | 66.4 |
| GRPO | 57.5 | 45.4 | 92.9 | 94.4 | 49.7 | 62.2 | 67.0 |
| DAPO | 61.0 | 46.0 | 92.8 | 94.1 | 49.8 | 64.8 | 68.0 |
| Dr.GRPO | 61.2 | 44.5 | 91.7 | 94.2 | 49.2 | 64.5 | 67.5 |
| RLOO | 59.8 | 47.9 | 91.2 | 93.9 | 48.6 | 62.3 | 67.3 |
| Method | LiveCodeBench | CodeForces | HumanEval+ | ||
|---|---|---|---|---|---|
| Avg@16 | Pass@16 | Rating | Percentile | Pass@16 | |
| Qwen3-4B | 30.5 | 40.9 | 578.8 | 1.2 | 89.0 |
| PPO | 39.7 (+9.2) | 57.7 (+16.8) | 1377.6 (+798.8) | 71.9 (+70.7) | 95.1 (+6.1) |
| GRPO | 39.5 (+9.0) | 55.1 (+14.2) | 1267.9 (+689.1) | 63.1 (+61.9) | 95.7 (+6.7) |
| DAPO | 41.0 (+10.5) | 52.3 (+11.4) | 1112.5 (+533.7) | 46.7 (+45.5) | 95.7 (+6.7) |
| Dr.GRPO | 41.6 (+11.1) | 57.7 (+16.8) | 1195.0 (+616.2) | 55.3 (+54.1) | 95.7 (+6.7) |
| Method | ALFWorld | WebShop | |||||||
| Pick | Look | Clean | Heat | Cool | Pick2 | All | Task Score | Succ. | |
| GPT-4o ( Hurst et al., 2024 ) | 75.3 | 60.8 | 31.2 | 56.7 | 21.6 | 49.8 | 48.0 | 31.8 | 23.7 |
| DeepSeek-V4-Pro | 69.0 | 66.7 | 41.9 | 26.3 | 52.2 | 60.0 | 51.6 | 77.9 | 33.3 |
| Prompting Backbone | 33.4 | 21.6 | 19.3 | 6.9 | 2.8 | 3.2 | 14.8 | 26.4 | 7.8 |
| Prompting ReAct | 48.5 | 35.4 | 34.3 | 13.2 | 18.2 | 17.6 | 31.2 | 46.2 | 19.5 |
| PPO | 92.3 | 64.0 | 92.5 | 89.5 | 80.3 | 68.8 | 80.4 | 81.4 | 68.7 |
| Method | AIME24 | AIME25 | AMC23 | MATH500 | Minerva | Olympiad | Average |
|---|---|---|---|---|---|---|---|
| GRPO | 57.5 | 45.4 | 92.9 | 94.4 | 49.7 | 62.2 | 67.0 |
| GRPO+Ent ( ) | 63.2 | 48.4 | 90.5 | 93.7 | 47.9 | 62.8 | 67.8 |
| GRPO+Ent ( ) | 57.1 | 47.7 | 91.9 | 92.8 | 47.0 | 63.8 | 66.7 |
| TAMPO | 63.3 | 47.1 | 92.0 | 95.2 | 49.2 | 61.0 | 68.0 |
| TGRL (ours) | 63.4 | 49.4 | 92.7 | 95.3 | 48.9 | 66.4 | 69.4 |
| Scale | Method | AIME24 | AIME25 | AMC23 | MATH500 | Minerva | Olympiad | Average |
|---|---|---|---|---|---|---|---|---|
| 14B | GRPO @ | 57.0 | 44.6 | 91.9 | 93.6 | 48.9 | 63.0 | 66.5 |
| GRPO @ | 57.5 | 45.4 | 92.9 | 94.4 | 49.7 | 62.2 | 67.0 | |
| TGRL (Uniform Token Credit) | 61.0 | 46.9 | 91.3 | 93.4 | 48.8 | 68.1 | 68.2 | |
| TGRL (ours) | 63.4 | 49.4 | 92.7 | 95.3 | 48.9 | 66.4 | 69.4 | |
| 32B | GRPO @ | 55.8 | 44.4 | 85.6 | 93.0 | 48.0 | 64.6 | 65.1 |
| GRPO @ | 60.4 | 44.8 | 88.9 | 92.2 | 48.5 | 62.1 | 66.2 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Selector | Count | Flip rate | Mean score drop | Mean position ratio | Mean JS |
|---|---|---|---|---|---|
| JS | 224 | 12.1% | 0.241 | 0.444 | 0.0427 |
| Low-Margin | 224 | 11.6% | 0.232 | 0.482 | 0.0309 |
| Matched-Random | 1120 | 8.7% | 0.173 | 0.458 | 0.0064 |
| Entropy | 224 | 7.6% | 0.152 | 0.468 | 0.0385 |
| Protocol | Method | Best AIME-Comb. | Time of Best (h) | Step | Remarks |
| (i) Post-warmup -hour fixed-budget comparison (TGRL vs. GRPO) | |||||
| TGRL vs. GRPO | TGRL (ours) | 1.0833 | 11.01 | 80 | h absolute |
| TGRL vs. GRPO | GRPO | 1.0354 | 8.86 | 100 | h absolute |
| (ii) From-start -hour fixed-budget comparison (TGRL vs. DAPO) | |||||
| TGRL vs. DAPO | TGRL (ours) | 1.0604 | 10.21 | 60 | from training start |
| TGRL vs. DAPO | DAPO | 0.9521 | 10.01 | 30 | from training start |
| Method | AIME24 | AIME25 | AMC23 | MATH500 | Minerva | Olympiad | Average |
|---|---|---|---|---|---|---|---|
| Qwen3-4B | 49.4 | 35.6 | 78.8 | 84.6 | 40.3 | 58.2 | 57.8 |
| GRPO | 54.2 | 39.2 | 83.9 | 90.8 | 44.4 | 64.2 | 62.8 |
| Dr.GRPO | 49.0 | 40.0 | 85.6 | 92.7 | 45.1 | 66.4 | 63.1 |
| RLOO | 53.1 | 39.0 | 85.8 | 81.8 | 40.4 | 57.6 | 59.6 |
| TAMPO | 54.0 | 39.4 | 85.9 | 92.2 | 44.2 | 64.2 | 63.3 |
| TGRL (ours) | 56.2 | 42.7 | 81.4 | 92.5 | 44.9 | 69.0 | 64.5 |
| Peak | Peak step | Final@150 | AUC@150 | Step to 0.94 | ||
|---|---|---|---|---|---|---|
| 0.8 | 0.5 | 0.808 | 130 | 0.758 | 0.698 | – |
| 1.0 | 0.7 | 0.942 | 140 | 0.900 | 0.822 | 140 |
| 1.2 | 0.9 | 0.963 | 70 | 0.933 | 0.885 | 70 |
| 1.4 | 1.1 | 0.923 | 50 | 0.900 | 0.846 | – |
| 1.6 | 1.3 | 0.975 | 140 | 0.948 | 0.814 | 110 |
| Method | AIME24 | AIME25 | AMC23 |
|---|---|---|---|
| Qwen3-14B | |||
| GRPO | |||
| TGRL (ours) | |||
| Qwen3-32B | |||
| GRPO | |||
| TGRL (ours) | |||
| Method | AIME24 | AIME25 | AMC23 | MATH500 | Minerva | Olympiad | Average |
|---|---|---|---|---|---|---|---|
| GRPO ( ) | 60.8 | 46.7 | 92.5 | 94.4 | 48.7 | 64.1 | 67.9 |
| TGRL ( , ours) | 66.0 | 49.4 | 92.2 | 93.7 | 48.9 | 64.3 | 69.1 |
| Split | AIME24 | AIME25 | AMC23 | MATH500 | Minerva | Olympiad | Average |
|---|---|---|---|---|---|---|---|
| 63.4 | 49.4 | 92.7 | 95.3 | 48.9 | 66.4 | 69.4 | |
| 60.4 | 48.8 | 87.8 | 93.0 | 48.4 | 64.3 | 67.1 | |
| 61.0 | 49.8 | 93.9 | 92.9 | 47.3 | 63.8 | 68.1 |
| Hyperparameter | Math (Qwen3-14B/32B) | Code (Qwen3-4B) |
|---|---|---|
| Training data | DeepScaleR | DeepCoder |
| Reasoning mode | Qwen3 think | Qwen3 think |
| Rollouts per prompt | 4 | 4 |
| Low-/high-temperature split | ||
| Low-temperature-group actor update | disabled | disabled |
| Warmup steps | 20 | 20 |
| Hyperparameter | ALFWorld | WebShop |
|---|---|---|
| Policy model | Qwen2.5-7B-Instruct | Qwen2.5-7B-Instruct |
| Engine | vLLM | vLLM |
| Training data size | 16 | 16 |
| Validation data size | 128 | 64 |
| Rollouts per prompt | 8 | 8 |
| Warmup steps | 20 | 25 |
| Control item | Protocol |
|---|---|
| Backbone and initialization | For each domain and scale, all methods start from the same policy model checkpoint: Qwen3 variants for math/code and Qwen2.5-7B-Instruct for agent experiments. No baseline is initialized from a stronger checkpoint than TGRL. |
| Training data and verifier | Within each domain, all methods use the same training data source, prompt pipeline, verifier/reward function, and environment interface. Math uses DeepScaleR, code uses DeepCoder, and agent experiments use the same ALFWorld and WebShop environments described in Table 14 . |
| Rollout budget | All math and code methods use the same total rollout budget responses per prompt. TGRL allocates this fixed budget into , while single-temperature baselines spend all four rollouts at their rollout temperature. Agent experiments use the matched budget reported in Table 14 . |
| Training rollout decoding | Single-temperature RLVR baselines use the same rollout decoding controls as the high-temperature branch unless the method defines an adaptive policy over temperature. In particular, top- , top- , maximum response length, and the rollout budget are matched. TAMPO is allowed to adapt temperature because this is its method-specific mechanism, but it receives no additional rollouts. |
| Evaluation decoding | All math and code checkpoints are evaluated with the same decoding configuration: , top- , top- off, and maximum response length . Agent evaluation uses the same environment limits and success metrics across methods. |
| Shared optimization settings | Optimizer family, learning-rate scale, batch size, prompt/response length limits, PPO mini-batch size, number of training epochs, and warmup convention are matched whenever they are not part of a baseline-specific algorithmic definition. Baseline-specific terms, such as entropy bonuses or decoupled clipping rules, are added on top of the shared training recipe rather than changing the data or evaluation protocol. |