An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
Figures & tables
Figure 1: Project overview . We investigate whether post-training fundamentally broadens the agentic reasoning capacity of its base model. (a) Our observations suggest that post-training sharpens the policy to improve sampling efficiency and consistency while compromising coverage. To quantify this systematically, (b) we present Sharpening Tax , measuring the difference in test-time scalability between base and post-trained LLMs. To lower the tax, (c) we then propose Posterior-Tempered Group Sampling (PTGS) , dynamically adjusting the sampling temperature based on task difficulty.
Figure 2: Base models catch up to their post-trained counterparts on agentic tasks . We draw pass@ K curves (up to 128 rollouts per task) for four different models on three benchmarks. Post-trained models (orange) dominate at small K , but pre-trained base models equipped with a light harness (teal) scale more steeply and cross over in most settings.
Figure 3: Larger models pull the crossover points earlier. Pass@ K test-time scaling curves of harness-equipped gemma-4 base models (teal) vs. their post-trained counterparts (orange) with model size increasing from left to right. Overall, the crossover budget at which the base model overtakes post-trained model shrinks as model scale grows.
Figure 4: Post-training bimodalizes per-task success rates ( gemma-4-31B , K=128 rollouts). The base model shows a spread distribution, whereas post-trained model polarizes it into a bimodal shape by pushing most of the intermediate pass given compute mass to the two extremes: always solved and never solved. See Appendix C.2 for full results.
Figure 5: Sharpening Tax across rollout budgets. TaxS(k) of eight models (the largest and smallest backbone of each family) across three benchmarks; positive values indicate the base model’s scaling advantage over RL. Shades represent 95% bootstrap confidence intervals. The tax is positive at almost every budget for the largest backbones, and for the smallest backbones it grows from negative values toward zero or above as the test-time budget scales.
Figure 6: Sharpening Tax extrapolates to larger budgets; predicts the consistency gap and the value of additional test-time compute. Each panel reports Spearman rank correlation ρ over 42 model-benchmark combinations, estimated and validated over disjoint task subsets. ( Left ) TaxS(8) vs. future tax TaxS(32) . ( Middle ) TaxS(8) estimated from 8 rollouts vs. consistency gap Δpass8 . ( Right ) Mean Spearman ρ of candidate predictors for other evaluation metrics.
Figure 7: Illustration of posterior-tempered group sampling . Given a policy, we model its success rate for each task (prompt) x as a dynamic difficulty estimate to adaptively determine rollout sampling temperature in RL training pipeline. It encourages exploratory behavior for hard prompts while inducing conservative behavior for easy ones.
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
#
Base checkpoint
Post-trained checkpoint
Family
1
google/gemma-4-E4B
google/gemma-4-E4B-it
Gemma-4
2
google/gemma-4-12B
google/gemma-4-12B-it
Gemma-4
3
google/gemma-4-26B-A4B
google/gemma-4-26B-A4B-it
Gemma-4
4
google/gemma-4-31B
google/gemma-4-31B-it
Gemma-4
5
mistralai/Ministral-3-3B-Base-2512
mistralai/Ministral-3-3B-Instruct-2512
Ministral-3
6
mistralai/Ministral-3-8B-Base-2512
mistralai/Ministral-3-8B-Instruct-2512
Ministral-3
Appendix
Table 1: The 14 base/post-trained checkpoint pairs used in our analysis , ranging from 3B to 35B parameters. gemma-4-26B-A4B and Qwen3.5-35B-A3B are mixture-of-experts models with about 4B and 3B active parameters, respectively.
Figure 9: Ablation on thinking mode. Coverage (pass@ k ) and consistency ( passk ) curves of the base and post-trained models of gemma-4-31B (left) and Qwen3.5-35B-A3B (middle) on BFCLv4 MT, with the thinking mode of the post-trained model turned off (our default setup) and on, and the corresponding TaxS(k) (right), as a function of the rollout budget k per task. Thinking mode improves the accuracy and coverage of the post-trained models, but it still doesn’t change the bold conclusion—base models still catch up at large k , and the tax remains non-negative.
Base model
Post-trained model
w/o harness
w/ harness
w/o harness
w/ harness
Backbone
Benchmark
p@1
p@32
p@1
p@32
Δ p@32
p@1
p@32
p@1
p@32
Δ p@32
gemma-4-E4B
BFCL
0.0
0.0
2.2
21.0
+21.0
23.7
38.0
12.0
21.0
-17.0
WebShop
1.4
16.8
0.9
10.2
-6.6
32.4
56.2
4.1
22.8
-33.4
ACEBench
23.8
77.6
30.8
77.7
+0.1
81.4
89.1
34.0
68.1
-20.9
gemma-4-31B
BFCL
0.0
0.0
32.8
86.5
+86.5
78.0
83.0
7.5
21.5
-61.5
Appendix
Table 2: Harness ablation for base and post-trained models. pass@ 1 and pass@ 32 (%) of eight backbones with (w/) and without (w/o) the harness, using N=32 rollouts per task and the same tasks, seed, temperature, and step budget in all four settings. Δ p@32 is the pass@ 32 gain from the harness (w/ minus w/o), and the larger value of each w/o–w/ pair is in bold. For base models, “w/o harness” serves the checkpoint with the chat template and native tool-call parser of its post-trained counterpart. For post-trained models, “w/ harness” sends the plain-text harness prompt through the completion endpoint, without the chat template or the native parser.
BFCLv4 multi-turn
WebShop
ACEBench
temperature
0.4
0.7
0.7
top_ p
0.95
0.95
0.95
max tokens per generation
1024
256
1200
context length
16384 / 32768
16384
16384 / 32768
rollouts per task ( N )
128
128
128
reported K grid of pass@ K
1,2,4,8,16,32,64,128
Appendix
Table 3: Decoding configuration for each benchmark , shared by the base and post-trained models of all 14 checkpoint pairs. Every model-benchmark combination is evaluated with N=128 rollouts per task.
Figure 10: Aggregate test-time scaling curves of the base and post-trained models for the largest (left) and smallest (right) backbones. Each curve is the mean pass@ k over 12 model-benchmark combinations (one backbone per family on each of the three benchmarks), and the bands show one standard deviation across combinations. For the largest backbones, the mean base curve overtakes the mean post-trained curve at k≈22 and ends clearly ahead at k=128 . For the smallest backbones, the base curve never catches up within the budget.
Figure 11: Aggregate coverage (pass@ k ) and consistency ( passk ) curves of the base and post-trained models for the largest (left) and smallest (right) backbones , averaged over 12 model-benchmark combinations (one backbone per family on each of the three benchmarks). The consistency of the base model drops to nearly zero as k grows, but that of the post-trained model stays much higher at both scales. For the largest backbones, more samples let the base model overtake the post-trained model in coverage, but not in consistency.
Figure 12: Aggregate per-task success-rate distributions of the base and post-trained models for the largest (left) and smallest (right) backbones , averaged over 12 model-benchmark combinations (one backbone per family on each of the three benchmarks). For the largest backbones, post-training increases both the never-solved share and the always-solved share (the latter from nearly 0% to 37%), i.e., it loses some solvable tasks and fully solves others. For the smallest backbones, post-training instead reduces the never-solved share by repairing the base model.
BFCLv4 MT
WebShop
ACEBench
Backbone
Task category
Base
Post
Base
Post
Base
Post
\Block 3-1 gemma-4-E4B
Always pass (any k )
0.0
12.5 ( ↑ )
0.0
11.4 ( ↑ )
0.0
66.4 ( ↑ )
Pass given compute
27.0
26.0 ( ↓ )
48.6
53.8 ( ↑ )
85.6
22.9 ( ↓ )
Always fail (any k )
73.0
61.5 ( ↓ )
51.4
34.8 ( ↓ )
14.4
10.8 ( ↓ )
\Block 3-1 gemma-4-12B
Always pass (any k )
0.0
60.5 ( ↑ )
0.0
14.0 ( ↑ )
0.5
83.2 ( ↑ )
Pass given compute
47.0
13.5 ( ↓ )
78.0
40.6 ( ↓ )
90.0
7.0 ( ↓ )
Appendix
Table 4: Task categories of all 14 base/post-trained pairs at K=128 . Each task is classified by its outcomes over 128 rollouts as always pass (every rollout succeeds), pass given compute (some but not all rollouts succeed), or always fail (no rollout succeeds). Numbers are percentages of tasks in each benchmark. Teal ( ↑ ) and purple ( ↓ ) mark an increase and a decrease after post-training. For the larger backbones of every family, the middle category shrinks and its tasks move to the two extremes. For the smallest backbones on BFCL, post-training can instead enlarge the middle category by moving tasks out of always fail, as it improves the tool-call formatting capability of the base model. This is the setting in which the tax can become negative.
Figure 13: Per-task success-rate distributions of the Gemma-4 base and post-trained models (empirical success rate over 128 rollouts per task) for every backbone and benchmark. Post-training moves most of the intermediate mass to the two extremes at every model size and on every benchmark. Among the four families, this polarization is the strongest for Gemma-4, which also has the largest mean tax.
Figure 14: Per-task success-rate distributions of the Qwen2.5 base and post-trained models (empirical success rate over 128 rollouts per task) for every backbone and benchmark. Post-training moves the intermediate mass to the two extremes, as in the other families.
Figure 15: Per-task success-rate distributions of the Qwen3.5 base and post-trained models (empirical success rate over 128 rollouts per task) for every backbone and benchmark. Qwen3.5 shows the softest sharpening among the four families.
Figure 16: Per-task success-rate distributions of the Ministral-3 base and post-trained models (empirical success rate over 128 rollouts per task) for every backbone and benchmark. The polarization grows with model scale. For the 3B backbone, the base model rarely succeeds on BFCL and WebShop and post-training largely improves it, while the 8B and 14B backbones show the shift of mass to the two extremes as we observed in other backbones.
Figure 17: Coverage (pass@ k ) and consistency ( passk ) curves of the Gemma-4 base and post-trained models for every backbone and benchmark. The two curves of the post-trained model stay close together, as expected from a bimodalized policy. In contrast, the consistency of the base model drops toward zero while its coverage keeps rising. The base model overtakes the post-trained model in coverage at a smaller budget for larger backbones.
Figure 18: Coverage (pass@ k ) and consistency ( passk ) curves of the Qwen2.5 base and post-trained models for every backbone and benchmark. The consistency of the base model drops toward zero while its coverage keeps rising, and the two curves of the post-trained model stay much closer together.
Figure 19: Coverage (pass@ k ) and consistency ( passk ) curves of the Qwen3.5 base and post-trained models for every backbone and benchmark. Since Qwen3.5 has the softest sharpening among the four families, the consistency curves of its post-trained models decay slightly instead of staying flat, and the coverage gaps between the base and post-trained models at k=128 are the smallest among the four families.
Figure 20: Coverage (pass@ k ) and consistency ( passk ) curves of the Ministral-3 base and post-trained models for every backbone and benchmark. The consistency of the base model drops toward zero while its coverage keeps rising, and the two curves of the post-trained model stay much closer together.
Figure 21: Coverage lost versus sampling efficiency gained by post-training at K=8 (left) and K=128 (right). The y -axis is the coverage lost, pass@Kbase−pass@KRL , and the x -axis is the efficiency gained, avg@KRL−avg@Kbase , where avg@ K is the mean success rate over K rollouts. Each point is one model–benchmark combination, colored by family, and the marker size grows with the total number of parameters. At K=8 , most points lie below zero, i.e., post-training is beneficial in both coverage and efficiency when compute is scarce. At K=128 , many points, mostly from the larger backbones, cross above zero, implying that post-training still gains efficiency, but now pays for it with lost coverage.
Figure 22: Per-task scaling curves on WebShop for the smallest and largest Gemma-4 backbones ( gemma-4-E4B , top two rows; gemma-4-31B , bottom two rows). In each block, the top row shows five base (teal) and five post-trained (orange) task curves from the harder (5–25%), intermediate (40–60%), and easier (75–95%) rank quantiles. Tasks are ranked separately for each policy by their success counts, among the tasks that at least one policy solves at least once. The bottom row shows the density of all 500 per-task curves of each policy and their mean pass@ k , with pass@ 1 and pass@ 128 given in the legend. All curves use the unbiased estimator from 128 rollouts, and a flat curve at zero means no observed success.
Figure 23: Per-task scaling curves on BFCLv4 for the smallest and largest Gemma-4 backbones ( gemma-4-E4B , top two rows; gemma-4-31B , bottom two rows). In each block, the top row shows five base (teal) and five post-trained (orange) task curves from the harder (5–25%), intermediate (40–60%), and easier (75–95%) rank quantiles. Tasks are ranked separately for each policy by their success counts, among the tasks that at least one policy solves at least once. The bottom row shows the density of all 200 per-task curves of each policy and their mean pass@ k , with pass@ 1 and pass@ 128 given in the legend. All curves use the unbiased estimator from 128 rollouts, and a flat curve at zero means no observed success.
Figure 24: Global temperature scaling trades single-sample accuracy for coverage rather than improving both. pass@ k of the post-trained model of the largest backbone in each family (columns) on BFCLv4 MT, WebShop, and ACEBench (rows), at the default temperature of each benchmark (0.4 for BFCL and 0.7 for the others) and at T∈{1.0,1.3} . Shaded bands show ±1 bootstrap standard error over tasks (200 resamples). A higher temperature usually improves pass@ k at larger budgets, but pass@ 1 drops in some cases.
Figure 25: Sharpening Tax as a function of the rollout budget for all 14 base/RL pairs. TaxA(k) (top) and TaxS(k) (bottom) on BFCLv4 MT, WebShop, and ACEBench, with one curve per checkpoint pair and shaded 95% bootstrap confidence intervals. Positive values mean that post-training reduces test-time scalability relative to the base model. The raw tax starts near zero and grows with the budget for almost every pair, while the calibrated tax can either rise or fall with the budget. Both variants stay negative at k=128 only for Ministral-3-3B , Qwen2.5-3B , and Qwen2.5-7B on BFCL, where post-training repairs the tool-call formatting of the base model.
Benchmark
Pair
CB
CP
ΔC [95% CI]
CB
CP
Tax A
Tax S
BFCLv4 MT
Gemma 4 E4B
0.270
0.385
-0.115 [-0.190,-0.050]
0.194
0.360
+6.59
+0.046
BFCLv4 MT
Gemma 4 12B
0.470
0.740
-0.270 [-0.345,-0.200]
0.380
0.733
+10.59
+0.075
BFCLv4 MT
Gemma 4 26B-A4B
0.625
0.810
-0.185 [-0.255,-0.120]
0.538
0.803
+10.18
+0.072
BFCLv4 MT
Gemma 4 31B
0.905
0.835
+0.070 [+0.015,+0.120]
0.838
0.831
+8.11
+0.073
BFCLv4 MT
Ministral 3 3B
0.240
0.615
-0.375 [-0.445,-0.300]
0.201
0.562
-1.71
-0.027
BFCLv4 MT
Ministral 3 8B
0.540
0.645
-0.105 [-0.190,-0.025]
0.454
0.611
+6.76
+0.046
Appendix
Table 5: Summary of the 42 model–benchmark evaluations at k=128 . CB and CP are the coverage ( pass@128 ) of the base and post-trained (RL) models, ΔC=CB−CP is the coverage gap with a 95% paired task-bootstrap confidence interval, and CB and CP are the budget-averaged coverage C(128)=1281∑k=1128pass@k of the two models. The raw scalability of each model is A(128)=128(C−C(128)) , and the last two columns are the raw and calibrated Sharpening Tax at k=128 . In every combination, the post-trained model reaches a larger fraction of its ceiling ( C/C ) than the base model.
Task categories (% of tasks)
Contributions to Tax A (128)
Benchmark
Pair
Lost
Gained
Both
Neither
Lost ( + )
Gained ( − )
Shape
Tax A (128)
BFCLv4 MT
Gemma 4 E4B
8.0
19.5
19.0
53.5
+3.88
-2.25
+4.97
+6.59
BFCLv4 MT
Gemma 4 12B
4.0
31.0
43.0
22.0
+1.73
-0.21
+9.07
+10.59
BFCLv4 MT
Gemma 4 26B-A4B
4.5
23.0
58.0
14.5
+0.82
-0.81
+10.17
+10.18
BFCLv4 MT
Gemma 4 31B
10.5
3.5
80.0
6.0
+2.43
-0.05
+5.73
+8.11
BFCLv4 MT
Ministral 3 3B
5.0
42.5
19.0
33.5
+1.73
-5.41
+1.96
-1.71
Appendix
Table 6: Decomposition of TaxA(128) by task movement for all 42 model-benchmark combinations. Each task is assigned to a category by which model solves it at least once in 128 rollouts, i.e., Lost (only the base model), Gained (only the post-trained model), Both , or Neither . The left block reports the proportion of tasks in each category. The right block splits TaxA(128) into the retry value lost on Lost tasks (positive), the retry value recovered on Gained tasks (negative), and the difference in curve shape on tasks solved by both models. The three contributions sum to the total in the last column.
Figure 26: Per-task success rate movement from the base to the post-trained model for the largest backbone of each family on BFCLv4 MT, WebShop, and ACEBench. Each point is one task, placed by its empirical success rate over 128 rollouts under the base model ( p^iB , x -axis) and the post-trained model ( p^iP , y -axis). Colors mark how each task moves, i.e., lost ( pB>0 , pP=0 ), gained ( pB=0 , pP>0 ), sharpened ( pP>pB>0 ), softened ( pB>pP>0 ), or unchanged, and the inset of each panel reports the proportion of each type among all tasks. The top row pools the four models, and each remaining row shows one model. Most tasks move far from the diagonal. Many sharpened tasks reach the top edge, while the lost tasks along the bottom edge show the coverage lost by post-training.
Figure 27: Mean scaling curves of the tasks with the highest and lowest tax contributions, averaged over all 14 backbones. For each benchmark and backbone, we select the five tasks with the highest (top row) and lowest (bottom row) signed contributions to TaxS(128) in Eq. ( 10 ). Each panel shows the mean pass@ k curve of the base (teal) and post-trained (orange) models over the selected tasks of all 14 backbones (70 tasks per panel).
Figure 28: Scaling curves of the tasks with the highest and lowest tax contributions for Gemma-4 (E4B, 12B, 26B-A4B, and 31B from left to right). For each benchmark (BFCL MT, WebShop, and ACEBench from top to bottom), the upper row shows the five tasks with the highest signed contributions to TaxS(128) (Eq. ( 10 )) and the lower row the five with the lowest, selected separately for each backbone. Each panel shows the pass@ k curves of the base (teal) and post-trained (orange) models on the same five tasks, computed from 128 rollouts per task without averaging.
Figure 29: Scaling curves of the tasks with the highest and lowest tax contributions for Ministral-3 (3B, 8B, and 14B from left to right). For each benchmark (BFCL MT, WebShop, and ACEBench from top to bottom), the upper row shows the five tasks with the highest signed contributions to TaxS(128) (Eq. ( 10 )) and the lower row the five with the lowest, selected separately for each backbone. Each panel shows the pass@ k curves of the base (teal) and post-trained (orange) models on the same five tasks, computed from 128 rollouts per task.
Figure 30: Scaling curves of the tasks with the highest and lowest tax contributions for Qwen2.5 (3B, 7B, 14B, and 32B from left to right). For each benchmark (BFCL MT, WebShop, and ACEBench from top to bottom), the upper row shows the five tasks with the highest signed contributions to TaxS(128) (Eq. ( 10 )) and the lower row the five with the lowest, selected separately for each backbone. Each panel shows the pass@ k curves of the base (teal) and post-trained (orange) models on the same five tasks, computed from 128 rollouts per task.
Figure 31: Scaling curves of the tasks with the highest and lowest tax contributions for Qwen3.5 (4B, 9B, and 35B-A3B from left to right). For each benchmark (BFCL MT, WebShop, and ACEBench from top to bottom), the upper row shows the five tasks with the highest signed contributions to TaxS(128) (Eq. ( 10 )) and the lower row the five with the lowest, selected separately for each backbone. Each panel shows the pass@ k curves of the base (teal) and post-trained (orange) models on the same five tasks, computed from 128 rollouts per task.
Figure 32: Tax-guided routing between the base and post-trained policies. Success rates averaged over the four largest (top) and the four smallest (bottom) backbones on BFCLv4 MT, WebShop, and ACEBench. Each task is served by one policy chosen by one of five rules, i.e., base only (teal), post-trained only (orange), uniform random routing (gray), tax-based routing (red), and optimal routing (navy). Tax-based routing chooses the policy for each task based on its TaxS(32) estimated from 32 early pilot rollouts. Optimal routing is an oracle selector that assigns each task to the better policy on the full 128 rollouts, which upper-bounds any router. Filled bars show pass@ 128 and dashed inset bars show pass@ 1 under the same assignments, both computed from all 128 rollouts of the selected policy including the pilot rollouts. Error bars show the bootstrap uncertainty of pass@ 128 (2,000 resampled tasks).
Figure 33: Training dynamics of PPO and PPO with PTGS on FrozenLake. The panels show, from left to right, training success, validation success, average entropy of the output token distribution (thin lines for raw values and thick lines for the moving average), and TaxS(128) of the checkpoints at steps 50, 100, 150, and 200. The two methods have nearly identical training success and end at the same validation success, but PTGS keeps a much higher entropy throughout training and pays a smaller tax at every checkpoint after step 50.
Figure 34: Training dynamics of PPO and PPO with PTGS on Sokoban. The panels show, from left to right, training success, validation success, average entropy of the output token distribution (thin lines for raw values and thick lines for the moving average), and TaxS(128) of the checkpoints at steps 50, 100, 150, and 200. PTGS reaches higher validation success at every checkpoint and keeps improving until step 200, while PPO drops after step 150. The entropy of PPO decays toward zero while that of PTGS stays high, and PTGS pays a smaller tax than PPO from step 100 on.
Figure 35: Ablation on the temperature spread τ of PTGS on FrozenLake (a) and Sokoban (b). We compare PPO with PTGS for τ∈{1.2,1.3,1.4,1.5} against the PPO baseline. In each block, the top row shows validation success, training success, and average entropy of the output token distribution over training, and the bottom row shows pass@ 1 , pass@ 128 , and TaxS(128) of the intermediate checkpoints, with dashed lines for the base model. A larger τ keeps a higher entropy in both environments. τ=1.5 achieves the best final pass@ 1 and pass@ 128 in both environments and the smallest tax on FrozenLake. On Sokoban, the smallest temperature τ=1.2 only slightly raises the entropy above PPO and pays a similar tax.
Sokoban
FrozenLake
Method
pass@ 1
pass@ 128
pass128
TaxS(128)
pass@ 1
pass@ 128
pass128
TaxS(128)
Qwen2.5-7B-Instruct (base)
20.7
76.6
0.0
–
26.3
89.1
0.0
–
PPO, T = 0.5
45.6 ± 7.0
62.8 ± 10.7
26.2 ± 8.6
0.073 ± 0.025
58.7 ± 15.0
76.9 ± 12.1
41.2 ± 26.5
0.004 ± 0.053
PPO, T = 0.7
40.2 ± 26.8
51.9 ± 33.6
30.9 ± 23.1
0.075 ± 0.049
63.7 ± 7.1
70.3 ± 4.9
54.7 ± 11.4
0.046 ± 0.032
PPO, T = 1.0 (default)
46.5 ± 14.7
55.0 ± 11.3
36.2 ± 13.1
0.094 ± 0.016
63.7 ± 5.9
74.1 ± 7.2
50.9 ± 10.1
0.039 ± 0.026
PPO, T = 1.3
53.6 ± 7.6
58.8 ± 10.6
47.2 ± 4.8
0.100 ± 0.022
60.9 ± 12.1
70.3 ± 12.0
47.5 ± 17.5
0.042 ± 0.026
Appendix
Table 7: Fixed rollout-temperature ablation of PPO, compared with PTGS . Each PPO row samples its training rollouts at a fixed temperature T , where T=1.0 (shaded) is the PPO baseline of the main text. Values are the mean over five runs ± the 95% confidence interval.
Sokoban
FrozenLake
Method
pass@ 1
pass@ 128
pass128
TaxS(128)
pass@ 1
pass@ 128
pass128
TaxS(128)
Qwen2.5-7B-Instruct (base)
20.7
76.6
0.0
–
26.3
89.1
0.0
–
PPO, ancestral (default)
46.5 ± 14.7
55.0 ± 11.3
36.2 ± 13.1
0.094 ± 0.016
63.7 ± 5.9
74.1 ± 7.2
50.9 ± 10.1
0.039 ± 0.026
PPO, top- p 0.9
42.5 ± 12.7
46.6 ± 15.2
33.1 ± 9.2
0.111 ± 0.012
66.4 ± 6.2
74.1 ± 7.1
52.8 ± 13.3
0.041 ± 0.039
PPO, top- p 0.95
47.3 ± 2.8
55.9 ± 11.4
36.6 ± 9.9
0.091 ± 0.054
61.9 ± 8.0
72.5 ± 4.5
46.2 ± 16.4
0.030 ± 0.027
PPO, top- k 20
45.5 ± 9.3
55.3 ± 14.2
37.5 ± 7.3
0.092 ± 0.028
61.3 ± 7.3
75.6 ± 13.2
40.0 ± 24.7
0.031 ± 0.034
Appendix
Table 8: Rollout decoding ablation of PPO, compared with PTGS . Each PPO row applies one truncation rule to the training rollouts at T=1.0 , and the shaded row is the PPO baseline with ancestral sampling (the same runs as T=1.0 in Table 7 ). Values are the mean over five runs ± the 95% confidence interval.
Figure 36: Inference-time PTGS versus global temperature scaling. pass@ 1 ( x -axis) and pass@ 128 ( y -axis) of the frozen Qwen2.5-7B-Instruct on 64 tasks of each environment, where the upper right is better. Squares show global temperatures T∈[0.1,1.5] , triangles show inference-time PTGS with τ∈[1.2,2.0] and Tref=0.5 , and the ring marks the default inference temperature T=0.5 . PTGS lies beyond the global-temperature curve in both environments.
Figure 44
Figure 39: PTGS reduces the proportion of zero-success rollout groups. Percentage of the rollout groups at each training step in which none of the 16 rollouts succeeds, i.e., groups that carry no success signal for the update, for PPO and PPO with PTGS. The proportion is measured before the reward-variance filter and averaged over Sokoban and FrozenLake. Thin lines show the mean over the ten runs of each method (five per environment) with a 9-step moving average as the thick line.
Figure 40: PTGS versus PPO with a globally higher temperature on Sokoban (top) and FrozenLake (bottom). We compare PPO at the default rollout temperature ( T=1.0 ), PPO with every rollout sampled at T=1.3 or T=1.5 (the highest temperature PTGS can assign), and PPO with PTGS ( τ=1.5 ). The left panels show the average entropy of the output token distribution during training, computed at the rollout temperature of each run, with thin lines for the mean over five runs and thick lines for a 9-step moving average. The middle and right panels show pass@ 1 and pass@ 128 of the intermediate checkpoints, averaged over five runs, and the dashed line marks the base model.
Figure 41: A Sokoban task that PPO loses and PTGS recovers. The above example case describes the task and the opening of a successful rollout of the base model, and the left panel shows the initial state. The right panel shows the number of successful rollouts out of 128 at step 200 in each of the five training runs of PPO and PPO with PTGS, and the dashed line marks the base model, which solves the task in 4 of 128 rollouts. PPO never solves the task in four of its five runs, but PPO w/ PTGS solves it in all five runs and masters it in two.