Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the policy to update without observing a terminal reward. We make critic-based credit assignment reliable through three design choices. First, we introduce local readiness which uses critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem. Second, when available, we provide the critic with a reference solution from a previous successful rollout. Third, we assign credit over action chunks of 10k tokens rather than individual tokens, giving the critic a more meaningful portion of the trajectory to evaluate. We train Qwen3-4B on FineProofs-RL using AC2 and evaluate on IMO-ProofBench. AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs. This gain comes from two sources, (1) AC2 requires 25% fewer training steps to reach this score, and (2) each step generates fewer tokens because the policy does not need to continue every trajectory to completion. Conceptually, we demonstrate that we can remove the need to roll out every trajectory to completion, opening up a large previously unexplored design space for LLM RL algorithms.
Figures & tables
Figure 1: Advantage estimation in GRPO and AC2. GRPO rolls every sample out from the root to a terminal reward. AC2 starts from a replayed prefix s , generates a short chunk ai per sample, and reads the critic at the chunk’s end, Vθπ(s⋅ai) , where ⋅ denotes concatenation. The same group also trains the critic. L is a classification loss that fits the critic’s prediction at the prefix to the group-mean value, its bootstrapped target ( Section 3.3 ). Right: mean score against decoding compute for AC2, GRPO, and Prefix GRPO, which applies GRPO to full-length continuations of replayed prefixes ( Section 4.1 ).
Figure 2: Critic readiness and accuracy for AC2. The first three panels show critic metrics over the course of the main AC2 run, shown as the blue line in Figure 1 . First: fraction of sampled problems that are ready. Second: critic error over all sampled prefixes, Vθπ(s)−meanj(vj) , where vj is the terminal reward or the critic’s value of the continuation. Third: critic error on audited groups of critic ready problems, Vθπ(s)−meanj(rj) , where every continuation runs to completion, so the target is a mean of terminal rewards only. In the first three panels, faint lines show per-step values, dark lines average these over the trailing five steps. Fourth: critic MAE with and without the reference solution in its prompt.
Figure 3: Component ablations evaluated by mean score versus Decoding FLOPs. Left: variants that perform comparably to AC2 . AC2 w/o Audit removes full-length audits with little effect on performance. AC2 w/o Group & Audit uses one short continuation per ready prefix and the starting-prefix value as its baseline, achieving comparable scores. AC2 w/ stale replay buffer refreshes its replay buffer only every ten steps and continues to improve. Right: variants that fall behind. With AC2 w/o Group & Audit & local readiness , learning stalls and scores decline; this run branches off AC2 at the checkpoint of step 20. AC2 w/ correct-only buffer retains trajectories receiving at least six of seven judge points in the replay buffer, following Setlur et al. (2026) , and worsens later in training. AC2 w/ 2k chunks uses 2,000-token chunks and falls behind after step 50. The corresponding training-step plots are in Appendix Figure 6 .
Figure 4: Group-mean values are more accurate than prefix predictions, and critic-based advantages correlate with GRPO advantages. At step 80, the left panel compares meani(vi) with meani(ri) for 256 prefixes; colour indicates the percentage of continuations scored by the critic. The middle panel compares the prefix value prediction Vθπ(s) with meani(ri) for the same prefixes. The right panel compares A^i with ri−meanj(rj) for 2,640 responses in 165 groups containing at least one critic prediction. Details are deffered to Section A.4 .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The observed learning-rate sweep. Panels show mean score against Decoding FLOPs and Steps.
Figure 6: Component ablations by steps. The panels show the comparisons from Figure 3 with training steps on the horizontal axis.
Figure 7: Chunk-size ablation with grouping and auditing enabled. AC2 uses 10,000-token chunks, AC2 w/ 2k chunks uses 2,000-token chunks and prefix-cut spacing.
Figure 8: Removing local readiness. AC2 w/o Group & Audit starts from the base model; the variant that also removes local readiness branches off AC2 at the checkpoint of step 20. The plot shows their sampled-problem ready fractions.
Figure 9: Best-of-16 scores for AC2 and GRPO.
Step
f
Groups
MAE
Signed bias
40
All
256
0.129
+0.112
40
f=0
94
0.000
+0.000
40
0<f<1
79
0.188
+0.169
40
f=1
83
0.218
+0.183
57
All
248
0.050
+0.008
57
f=0
93
0.000
+0.000
Appendix
Table 1: Error between meani(vi) and meani(ri) by f . Each observation is a group of 16 continuations. Bias is meani(vi)−meani(ri) , averaged over groups.
Figure 10: Critic value at the prefix against the mean terminal reward. Each point is one prefix s , with Vθπ(s) on the vertical axis and meani(ri) over its 16 continuations on the horizontal axis. Left: step 57 (248 prefixes). Right: step 160 (256 prefixes). Step 80 is shown in Figure 4 , middle.
Figure 11: Group mean of critic values against the group mean of terminal rewards at step 80. For the 157 groups with at least two continuations where vi=Vθπ(s⋅ci) , the mean of these vi (vertical axis) against the mean of the same continuations’ ri (horizontal axis).
Figure 12: Group mean of vi against the group mean of ri . Each point is one group of 16 continuations, with meani(vi) on the vertical axis and meani(ri) on the horizontal axis. Left: step 40. Right: step 57. Colour shows 100f , the percentage of the group’s continuations with vi=Vθπ(s⋅ci) .
Figure 13: Critic values and group means at steps 80 and 160. Left: at step 80, Vθπ(s) against meani(vi) for 256 prefixes. Right: at step 160, meani(vi) against meani(ri) for 256 groups, with colour showing 100f .
Figure 14: Value diagnostics at step 40. Left: for groups with at least two continuations where vi=Vθπ(s⋅ci) , the mean of these vi against the mean of the same continuations’ ri . Right: the critic-based advantage vi−meanj(vj) against the GRPO advantage ri−meanj(rj) , for groups with f>0 .
Figure 15: Value diagnostics at step 57. Left: for groups with at least two continuations where vi=Vθπ(s⋅ci) , the mean of these vi against the mean of the same continuations’ ri . Right: the critic-based advantage vi−meanj(vj) against the GRPO advantage ri−meanj(rj) , for groups with f>0 .
Figure 16: Value diagnostics at step 160. Left: for groups with at least two continuations where vi=Vθπ(s⋅ci) , the mean of these vi against the mean of the same continuations’ ri . Right: the critic-based advantage vi−meanj(vj) against the GRPO advantage ri−meanj(rj) , for groups with f>0 .
Figure 17: Critic value at the end of the action chunk against the terminal reward. Each point is one continuation with vi=Vθπ(s⋅ci) , showing Vθπ(s⋅ci) against ri . Left: step 40. Right: step 57.
Figure 18: Critic value at the end of the action chunk against the terminal reward. Each point is one continuation with vi=Vθπ(s⋅ci) , showing Vθπ(s⋅ci) against ri . Left: step 80. Right: step 160.
Step
Group n
MAE
Centered n
MAE
Individual n
40
153
0.312
2592
0.177
1919
57
144
0.121
2480
0.141
1685
80
157
0.169
2640
0.150
1848
160
153
0.159
2672
0.146
1892
Appendix
Table 2: Sample counts and MAEs for the value diagnostics. Group n counts groups with at least two continuations where vi=Vθπ(s⋅ci) . Centered n and individual n count continuations.
Configuration
Length source
Last cost step
AC2
Joint lengths
202
GRPO, 10−6
Step means
70
GRPO, 2×10−6
Step means
180
GRPO, 4×10−6
Step means
100
Prefix GRPO
Means + joint lengths
120
No audit, branch at 40
Joint lengths
123
Appendix
Table 3: Available decoding-cost records. Joint lengths retain per-request prefix/decode moments. Step means yield estimated costs. The last cost-data step can exceed the last evaluation.
Figure 19: GPU-hours against Decoding FLOPs per training step. Each point is one training step of the main AC2 run. The line is the least-squares fit.
Parameter
Symbol
Value
Group size
g
16 continuations per prefix
Audit fraction
α
1/4 ; count of ready problems rounded up
Replayed prefixes per step
nbatch
192
Fresh problems per step
nrefill
192 problems, one rollout each
Replay buffer capacity
—
256 trajectories
Total response budget
—
50,000 tokens
Appendix
Table 4: Main AC2 configuration and evaluation settings. Symbols follow Section 3 . The total response budget includes replayed prefixes, with b counting new tokens. Settings without a symbol in Section 3 are listed by name.
Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit. This characterization provides a unified basis for explaining phenomena across existing algorithms and guides the development of an improved actor-critic training procedure. Through this lens, an ideal teacher in On-Policy Distillation (OPD) acts as an implicit critic, yielding an expected policy gradient proportional to that induced by token-level credit. Response-level REINFORCE Leave-One-Out (RLOO) signals match the expected policy-gradient contribution of token-level credit despite their coarser granularity. We further establish approximate credit sparsity under bounded outcome rewards and show how intermediate critic errors in Generalized Advantage Estimation (GAE) can become comparable to the underlying credit. These motivate Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy. In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively. On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.
Credit assignment is a central challenge in reinforcement learning (RL). Classical actor-critic methods address this challenge through fine-grained advantage estimation based on a learned value function. However, learned value models are often avoided in modern large language model (LLM) RL because conventional discriminative critics are difficult to train reliably. We revisit value modeling and argue that this difficulty is partly due to limited expressiveness. In particular, representation complexity theory suggests that value functions can be hard to approximate under the one-shot prediction paradigm used by existing value models, and our scaling experiments show that such critics do not improve reliably with scale. Motivated by this observation, we propose Generative Actor-Critic (GenAC), which replaces one-shot scalar value prediction with a generative critic that performs chain-of-thought reasoning before producing a value estimate. We further introduce In-Context Conditioning, which helps the critic remain calibrated to the current actor throughout training. GenAC improves value approximation, ranking reliability, and out-of-distribution generalization, and these gains translate into stronger downstream RL performance than both value-based and value-free baselines. Overall, our results suggest that stronger value modeling is a promising direction for improving credit assignment in LLM reinforcement learning.
Zikang Shan, Han Zhong, Liwei Wang +1
Peking University · Microsoft Research Asia · Shanghai Jiao Tong University
A commonly accepted explanation of critic-free RL for LLMs, based on sequence-level rewards, is that it reinforces successful rollouts with a positive advantage while penalizing failed ones. In contrast, we study critic-free RL from a token-level perspective, revealing the token-flipping phenomenon: positive and negative rollouts exhibit remarkably similar proportions of tokens whose probabilities are boosted or suppressed during RL training. To explain this phenomenon, we further show that a token's change in probability is not fully determined by its own advantage; coupled gradient interactions with other tokens also play a non-negligible role. Specifically, these token coupling effects occur primarily between identical tokens that are both predicted with low confidence. Building upon this analysis, we propose the cancellation hypothesis: as a result of coupling, opposing signals cancel out for tokens shared by positive and negative rollouts, while tokens more specific to successful rollouts receive stronger reinforcement, thereby inducing hidden token-level credit assignment from rollout-level rewards. We support this hypothesis with complementary empirical evidence. (1) Compared with training on only positive rollouts, critic-free RL shifts updates from template and formatting tokens toward reasoning tokens; (2) Tokens boosted by critic-free RL consistently demonstrate higher value than suppressed tokens, regardless of whether they originate from positive or negative rollouts. Guided by this view, we implement two batching interventions to encourage or preserve cancellation in critic-free RL training: query-preserved mini-batching and reward-balanced batching. Despite their simplicity, these interventions improve RLVR training across multiple model scales, supporting cancellation as both an explanatory principle and a practical design criterion for critic-free RL training.
Tianhao Cheng, Zeyu Huang, Zihan Qiu +5
1Fudan University · 3The University of Edinburgh · 6Qwen Team, Alibaba Group +4