Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.
Figures & tables
Arm
Training reward
p
Steps
Used for
100% labels (PPO)
verifier
1.0
300
label efficiency
RFPO, 50% labels
verifier or Eq. ( 2 )
0.5
330
label efficiency
RFPO, 0% labels ( 3× )
Eq. ( 2 )
0
310/324/378
main result
Raw value
v , no calibration
0
256
App. F
Threshold only
1[v>0.6]
0
100
App. F
Continuous v ( 4× )
v−b(ℓ) , no threshold
0
87/141/220/257
App. F
Table 1: Training runs. Shaded rows share the initial policy and the initial critic and differ only in the reward; rows below the rule are RFPO variants with parts of the calibration removed or changed, followed by the 4,096 -token-cap pair added later (§ 5 ). Steps are recorded run lengths, which can exceed the 300-step comparison window. All runs train under a 5,120 -token cap unless the arm names another; the raw-value run predates that protocol. Appendix H maps each arm to its log.
Model
AIME25
AIME26
AMC23
GPQA
Avg
Qwen3-4B-Base †
3.1
3.5
25.9
33.0
16.4
+ SFT (45k CoT)
22.3
20.8
63.4
35.7
35.5
Initial policy πθ0
22.7
21.0
71.4
44.9
40.0
+ PPO, 100% labels
25.2
21.7
73.6
46.6
41.8
+ RFPO, 50% labels
22.7
21.2
72.8
48.5
41.3
+ RFPO, 0% labels
25.4
21.7
70.0
47.3
41.1
Table 2: Separate evaluation after 300 RL steps. Pass@1 (%) at 16 samples, temperature 0.7 and a 12,288 -token budget; Avg is the mean over the four suites. Each RL row is one checkpoint; the zero-label row covers one of the three runs, the other two being assessed through training-time validation. † Base is evaluated as plain completion, without the think template.
AIME 2026
AMC 2023
Arm
Steps
pass@1
pass@8
pass@1
pass@8
Initial policy πθ0
0
21.2
32.1
73.6
89.5
PPO, 100% labels
300
25.0 (190) / 21.7
41.7 (250) / 34.0
78.1 (120) / 71.9
93.2 (210) / 89.0
RFPO, 50% labels
330
25.8 (80) / 22.5
42.9 (300) / 39.3
78.1 (290) / 71.2
94.1 (290) / 89.0
RFPO, 0% labels, run 1
310
24.6 (220) / 20.0
42.7 (220) / 33.0
76.2 (290) / 73.1
93.4 (290) / 85.7
RFPO, 0% labels, run 2
324
26.2 (30) / 23.8
43.7 (130) / 39.3
75.9 (40) / 72.2
93.8 (180) / 90.9
Table 3: Peak and final validation across every run. Each cell is peak (step) / final over each run’s recorded length. The shaded rows summarize the three zero-label runs. All rows train under the 5,120 -token cap except the last two, the 4,096 -token-cap pair added later (Table 4 ).
AIME 2026
AMC 2023
Arm
pass@1
pass@8
pass@1
pass@8
GPU-h
RFPO, 0% labels, 4,096 cap
21.6 / 24.6
36.1 / 44.4
73.5 / 76.9
90.1 / 93.6
264
PPO, 100% labels, 4,096 cap
20.9 / 24.6
33.5 / 40.3
72.4 / 75.6
89.5 / 93.0
327
PPO, 100% labels, 5,120 cap
22.1 / 25.0
34.5 / 41.7
74.4 / 78.1
89.8 / 93.2
388
Table 4: 4,096-token cap, per benchmark. Validation at 12,288 tokens (%), each cell the mean over the evaluations at steps 10–300 / the peak in that window; GPU-hours for 300 steps at the median step time. The first two rows ran on the same node; the last row is supervised PPO at the full 5,120 -token cap (Table 3 ), for reference.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
pass@1 (%)
Model
AIME 2025
AIME 2026
AMC 2023
GPQA-D
Mean length
Cap hit (%)
Qwen3-4B-Base †
3.1
3.5
25.9
33.0
950
3.5
+ SFT
22.3
20.8
63.4
35.7
8,719
48.5
+ critic pretraining ( πθ0 )
22.7
21.0
71.4
44.9
3,793
2.3
Appendix
Table 5: What the supervised stages install. n=16 samples at temperature 0.7 with a 12,288 -token budget; mean length and cap hit are averaged over the four suites. SFT buys accuracy with runaway generation, and the PPO run that trains the critic brings length back down while adding accuracy, most on GPQA-Diamond. † Evaluated without the chat template and without thinking mode.
Setting
Critic pretraining, ph. 1
Phase 2
RL stage (PPO and RFPO)
Actor init
SFT
phase-1 step 500
πθ0 (cumulative 700)
Critic init
SFT, new value head
phase-1 step 500
Vˉϕ (cumulative 800)
Critic
trained
trained
PPO: trained; RFPO: frozen
Reward
verifier
verifier
PPO: verifier; RFPO: Eq. ( 2 )
Steps
500
300
300 compared (runs to 378)
Prompts × responses
64×8
128×8
128×8
Appendix
Table 6: Hyperparameters. Rows above the rule change between stages; rows below are shared by every stage.
Decile
1
2
3
4
5
6
7
8
9
10
Upper edge
1,735
2,137
2,455
2,782
3,150
3,556
3,884
4,469
5,079
5,120
b(ℓ)
+0.071
+0.087
+0.100
+0.065
+0.083
+0.063
+0.024
−0.070
−0.104
−0.224
Appendix
Table 7: Deployed length baseline. Lengths outside the fitted range take the nearest decile’s offset. The last decile holds the rollouts that reach the cap.
AIME 2026
AMC 2023
Arm
Window
pass@1
pass@8
pass@1
pass@8
Initial policy πθ0
step 0
21.2
32.1
73.6
89.5
PPO, 100% labels
210–300
22.5
34.4
74.7
90.4
RFPO, 50% labels
240–330
23.1
39.1
74.5
90.9
RFPO, 0% labels, run 1
220–310
22.6
37.3
73.3
89.9
RFPO, 0% labels, run 2
230–320
21.8
37.0
71.2
89.2
Appendix
Table 8: Late-training validation. Accuracy (%) averaged over each run’s last ten evaluations, the complement to the peaks of Table 3 . All rows train under the 5,120 -token cap except the last two, the 4,096 -token-cap pair added later (Appendix E.5 ).
maj@16 / pass@16 (%)
median tokens / truncated (%)
Model
AIME25
AIME26
AMC23
GPQA
AIME26
GPQA
+ SFT
20.0 / 50.0
20.0 / 43.3
60.0 / 92.5
26.3 / 89.4
12,288 / 63.5
10,197 / 40.7
Initial policy πθ0
20.0 / 40.0
20.0 / 33.3
72.5 / 92.5
37.9 / 87.9
3,963 / 4.0
2,900 / 2.1
PPO, 100% labels
23.3 / 50.0
23.3 / 50.0
70.0 / 95.0
40.4 / 86.9
4,151 / 0.8
2,923 / 0.4
RFPO, 50% labels
20.0 / 46.7
20.0 / 43.3
75.0 / 95.0
43.9 / 89.9
4,586 / 1.0
2,909 / 1.0
RFPO, 0% labels
23.3 / 46.7
20.0 / 40.0
70.0 / 90.0
41.9 / 89.9
4,920 / 1.2
3,359 / 1.3
Appendix
Table 9: Extended metrics of the step-300 checkpoints. Same evaluation as Table 2 : 16 samples, temperature 0.7, 12,288 -token budget.
Within-problem AUC
Critic step
AUC
all
truncated
finished
EV
τ∗
Precision
v∈/[0,1]
100
0.913
0.765
0.800
0.608
0.467
0.956
0.532
0.443
200
0.923
0.776
0.794
0.640
0.460
0.462
0.796
0.044
300
0.930
0.789
0.802
0.667
0.521
1.082
0.540
0.556
400
0.933
0.797
0.820
0.678
0.568
0.809
0.683
0.325
500
0.937
0.796
0.808
0.683
0.593
0.757
0.673
0.295
Appendix
Table 10: Measurement plane, per critic. Each row averages ten actor distributions of 16,384 rollouts. The shaded row is the frozen critic. Truncated rollouts are 19–54% of a cell depending on the actor.
Prefix (tokens)
AUC
Within
Within, no answer yet (problems)
Between-problem share
Finished
256
0.958
0.526
0.521 (46)
0.996
0.003
512
0.961
0.477
0.470 (46)
0.994
0.003
1,024
0.967
0.600
0.603 (46)
0.982
0.004
2,048
0.967
0.647
0.645 (46)
0.955
0.141
3,072
0.969
0.735
0.700 (41)
0.924
0.395
4,096
0.976
0.836
0.844 (30)
0.905
0.615
Appendix
Table 11: Prefix scan. “No answer yet” restricts to rollouts still running at that prefix length; “between-problem share” is the fraction of score variance explained by problem identity; “finished” is the fraction of rollouts already complete.
Generation cap (tokens)
1,024
2,048
5,120
8,192
12,288
Finished, real cap
0.007
0.179
0.735
0.941
0.967
EV, real cap
−0.492
+0.496
+0.701
+0.669
+0.663
EV, truncation proxy
+0.554
+0.592
+0.649
+0.664
—
Precision, real cap
0.213
0.675
0.895
0.896
0.894
Precision, truncation proxy
0.879
0.919
0.905
0.897
—
Appendix
Table 12: Real caps versus truncation. 16,384 rollouts per column; the proxy truncates 12,288 -token rollouts to the cap.
Arm
Strength
Steps
AIME peak / final
Macro peak
Δ length
Truncated
No debiasing
0%
87
23.8 / 17.1
48.8
−35.2%
0.05
Debiased up to 2,828
16%
141
24.2 / 22.1
49.3
−4.2%
0.23
Debiased up to 3,558
34%
220
25.4 / 19.6
51.1
+12.8%
0.26
Full debiasing curve
56%
257
29.2 / 16.2
53.3
+33.6%
0.97
56%, group-relative
56%
161
28.3 / 15.8
52.4
+33.4%
0.97
Appendix
Table 13: Continuous-score arms. Accuracies are validation pass@1 in points; length change compares mean training response length at the last and first step; “truncated” is the fraction of training rollouts at the 5,120 cap over the last five steps; “up to” means the debiasing term is held constant beyond that length. The arms start at 20.0–24.2 on AIME and 46.7–48.0 macro.
56% arm
Step 1
Peak (step 110)
Step 250
Restatements per response
1.79
4.14
67.94
Boxed answers per response
1.63
1.82
34.92
Never closes its reasoning block
0.250
0.363
0.914
Verbatim repetition (60-character shingles)
0.000
0.000
0.010
Mean response length (tokens)
3,812
4,556
5,107
Train-batch accuracy, full text
0.430
0.527
0.496
Appendix
Table 14: What the 56% arm’s decline is made of. 256 inspected rollouts per checkpoint; length, score gap and entropy are from the training log. Neither entropy nor the score gap registers the change.
Per-step component (median seconds)
PPO, 100% labels
RFPO, 0% labels
Generation
181
176
Old log-probabilities (forward)
43
42
Values (forward; also the reward in RFPO)
38
37
Advantages and weight sync
6
6
Actor update
160
156
Critic update
151
—
Appendix
Table 15: Measured step time and memory. Medians over steps 1–300; validation and checkpointing are excluded.
Run
Log
Used in
SFT
sft_4n_4581861
App. A
Critic pretraining, phase 1
h061_warmstart_phase1_ppo
§ 3.1 , App. A , E
Critic pretraining, phase 2
h065_warmstart_phase2_ppo
§ 3.1 , App. A , E
PPO, 100% labels
h066_label100_ppo
§ 5 , App. C – E
RFPO, 50% labels
h054_label50_car
§ 5 , App. C – E
RFPO, 0% labels, run 1
h056_label0_seed2_car
§ 5 , App. C – E
Appendix
Table 16: Run roster. Every RL run starts from the initial policy and frozen critic extracted from the phase-2 run, except the raw-score run, which used an earlier actor–critic pair and an 8,192 -token cap.
Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training. Despite their simplicity, existing critic-free approaches propagate a trajectory-level learning signal uniformly across all tokens in a trajectory. This requires full-trajectory policy updates for every rollout, leading to substantial optimization cost for long reasoning traces, even though intermediate prefixes often contain enough information to largely determine the final outcome. We propose Prefix-Sampling Proximal Policy Optimization (PS-PPO), a compute-efficient critic-free method for RLHF that exploits this temporal redundancy. PS-PPO introduces a prompt-conditioned cutoff distribution and samples a cutoff timestep for each trajectory. During the update pass, PS-PPO backpropagates only through the sampled prefix of each trajectory and applies an importance-weighting correction so that the resulting truncated gradient estimator remains unbiased with respect to the full-trajectory objective. Experiments on mathematical reasoning and RLHF benchmarks show that PS-PPO achieves large reductions in training compute and peak GPU memory, while maintaining accuracy comparable to strong critic-free baselines.
Doo Hwan Hwang, Kee-Eung Kim
Kim Jaechul Graduate School of AI, KAIST, Daejeon, South Korea.
Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO work as well as they do, and when they should provide the largest gains. We develop a value-gradient perspective of critic-free RL for LLM post-training. First, under a differentiable rollout and additive-noise parameterization, we show that the actor update is value-gradient-like in expectation: the backward pass propagates costates whose conditional expectation equals the value gradient. Second, for discrete transformer policies, we show that autodifferentiation through attention produces empirical costates that approximate this value signal, with an error controlled by the sampling gap and policy entropy. These results motivate a decomposition of RL impact into value gradient signal and reachable reward headroom, yielding a criterion for when RL should be most effective along a pretraining trajectory.
The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from scratch and applying RL, SFT, and SFT followed by RL directly to intermediate pre-training checkpoints. We find that RL is effective very early, and often matches the full SFT→RL pipeline early as well. Through experiments on harder problems, we find that targeted pre-training data composition is a strong lever for RL effectiveness, even more so than model scale. Beyond reasoning accuracy, applying RL directly to base checkpoints expands the model's distribution; the sharpening effect reported in recent work arises only when RL follows SFT. The general capabilities of the model remain essentially unchanged by RL, while they degrade following SFT. Finally, we merge RL and SFT objectives by parallel averaging, which outperforms across all other training methods discussed, across metrics, while preserving general capabilities. Together, these results suggest that LLM training might benefit from an expanded use of RL.