Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.
Figures & tables
Figure 1: SCAPO achieves the highest AIME accuracy on the two Qwen3 base models. Aggregate AIME 2024–2026 accuracy for SCAPO, GRPO, and FIPO. SCAPO reaches GRPO’s final accuracy in fewer than half as many policy optimization steps and achieves higher final accuracy than baselines at both model scales. Shaded regions mark SCAPO’s semifactual credit augmentation phase.
Figure 2: Semifactual instability varies across tokens. (a) Probability transitions under semifactual prompt interventions. (b) Distribution of mean token drift. (c) Relative category mean drift normalized by the overall mean, with reflection markers and discourse connectives showing higher sensitivity. Results use frozen Qwen3-4B-Base before RL.
Figure 3: Semifactual filtering improves rollout accuracy. Suppressing high-drift token candidates during rollout generation improves Qwen3-4B-Base’s accuracy by 14.2 percentage points on the same 1,000-question diagnostic panel.
Figure 4: Overview of SCAPO. (a) Semifactual prompt interventions are paired with a fixed group of sampled responses. (b) Teacher forcing measures sampled-token probability drift under each intervention. (c) Group-wise normalization and aggregation yield relative stability scores, retaining only their negative part. (d) The scaled stability correction, with gradients stopped, is added to the GRPO advantage to produce token-level advantages for policy optimization.
Benchmark
Base
GRPO
GSPO
SAPO
CF-GRPO
FIPO
SCAPO
Qwen3-4B-Base
In-Distribution
AIME 24–26
7.64
22.08
23.47
24.03
22.64
25.42
27.71
AMC 23–25
37.40
59.74
64.71
65.45
62.06
64.57
65.65
HMMT 25–26
2.08
11.71
13.69
15.18
14.09
15.28
16.17
BRUMO 25
18.54
28.96
33.75
33.54
32.08
35.00
37.71
SMT 25
12.74
27.12
26.06
27.12
26.77
28.18
30.31
Table 1: Overall accuracy on competition-level mathematical reasoning and out-of-distribution benchmarks (%). Base denotes the model before RL. AIME variants aggregate 2024–2026, while HMMT and AMC aggregate 2025–2026 and 2023–2025, respectively. Year-specific results are provided in Appendix F.1 . Bold and underline indicate the best and second-best results, respectively.
Method
AIME 24–26
AMC 23–25
HMMT 25–26
BRUMO 25
SMT 25
Avg.
GRPO
7.71
32.97
1.98
13.54
6.96
12.63
SCAPO
11.88
41.39
4.96
19.17
10.85
17.65
Credit signal source
Random shuffle
8.33
32.78
2.28
15.42
9.43
13.65
Counterfactual
6.32
30.51
1.69
12.92
6.96
11.68
Correction sign
Table 2: Ablations on Qwen3-1.7B-Base. SCAPO uses negative-only corrections.
Figure 5: Training time breakdown of SCAPO on Qwen3-4B-Base. Other includes validation, reward computation, old policy log-probability computation, and miscellaneous overhead.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Classification rule
Token examples
Share (%)
Mean dt
Relative mean
Reflection marker a
Matches the reflection lexicon after lexical normalization.
any , check , Now , However , verify , again
0.29
0.04496
2.74
Discourse connective
Matches the 80-form DiMLex-Eng lexicon after normalization. Excludes reflection markers.
and , for , if , or , Thus , Therefore
2.68
0.03807
2.32
Word
Contains a Unicode letter and matches neither lexical category nor the math-symbol rule.
the , of , is , to , we , in
40.92
0.02506
1.53
Punctuation
All non-whitespace characters have Unicode category P* . Excludes math symbols.
, . : # ), :\n
6.75
0.01956
1.19
Math symbol
Begins with \ or contains only characters from the fixed math/ L a T e X symbol set. Excludes numbers.
\ , = , + , - , { , }
26.62
0.00768
0.47
Number
Matches a signed integer, decimal, or percentage expression.
0 , 1 , 2 , 3 , 4 , 5
15.31
0.00612
0.37
Appendix
Table 4: Token categories and drift statistics underlying Figure 2 (c). Share is the percentage of the 870,586 response token positions assigned to each category. Mean is the category’s average drift dt . Relative mean is this average divided by the overall mean (0.01640). denotes a space and \n a newline.
Figure 6: Token-level semifactual sensitivity in a compound-interest problem.
Figure 7: Token-level semifactual sensitivity in a numeral-base problem.
Metric
Base sampling
d -filtered
Correct answers
158
300
Accuracy (%)
15.80
30.00
Paired correctness outcomes (question counts)
Filtered correct
Filtered incorrect
Baseline correct
98
60
Baseline incorrect
202
640
Appendix
Table 5: Decoding results on 1,000 questions with frozen Qwen3-4B-Base. The lower block reports question counts for the four combinations of baseline and filtered answer correctness.
Hyperparameter
Qwen3-4B-Base
Qwen3-1.7B-Base
Data and rollout settings
Training dataset
DAPO-Math-17K ( Yu et al., 2025 )
Maximum prompt length
1,024
Maximum response length
16,384
Rollout batch size (prompts)
128
Responses per prompt ( G )
8
Appendix
Table 6: Shared and method-specific training hyperparameters. Batch sizes count prompts unless stated otherwise. Each rollout step contains two policy optimization steps. Clipping offsets are relative to 1.
Benchmark
Problems
Samples per problem
AIME 2024 / 2025 / 2026
30 / 30 / 30
16
AMC 2023 / 2024 / 2025
40 / 45 / 42
16
HMMT Feb. 2025 / 2026
30 / 33
16
BRUMO 2025
30
16
SMT 2025
53
16
Omni-Math
2,821
1
Appendix
Table 7: Evaluation benchmarks and sample counts. Slash-separated problem counts follow the listed years. ∗ For GPQA-Diamond, we enumerate all 4!=24 permutations of the multiple-choice options to avoid contamination.
In-Distribution
Out-of-Distribution
Method
AIME
AMC
HMMT
NoOp-AIME
ThinkBench-AIME
24
25
26
23
24
25
25
26
24
25
26
24
25
26
Qwen3-4B-Base
Base
8.33
7.92
6.67
47.66
32.50
32.89
1.04
3.03
8.33
7.29
4.58
8.96
3.75
6.88
GRPO
23.54
23.13
19.58
66.41
55.69
57.74
10.83
12.50
21.25
21.25
19.17
25.42
22.92
19.58
GSPO
27.08
22.92
20.42
73.59
61.39
59.82
10.42
16.67
19.58
21.46
19.38
24.58
22.71
18.96
Appendix
Table 8: Yearly accuracy (%) for Table 1 . Bold and underlining denote the best and second-best results.
Figure 8: Pass@ k on the AIME 2024–2026 and AMC 2023–2025 benchmarks at both model scales.
Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verifies whether a trajectory succeeds but provides no direct feedback on the reasoning path that produced it. Before success, prerequisite progress on hard problems receives no reward signal; after success, outcome rewards cannot distinguish well-organized correct trajectories from redundant or locally flawed ones. We introduce SCOPE-RL (Scaffolded Chain Optimization with Process Efficiency), a two-stage framework that densifies this anchor while retaining the GRPO update: Adaptive Scaffolded RL adds prefix-decomposed verifiable rewards on answer-hidden sub-question chains before success, and Quality-Aware Process RL applies correctness-gated process-shape rewards to refine correct trajectories after success. An expert-validated Step-Quality Evaluation Protocol evaluates useful-step density, error localization, and token efficiency beyond final-answer accuracy. On Qwen3-8B-Instruct trained on DAPO-Math and Big-Math, SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO; the gains hold under GSPO and on Qwen3-0.6B-Instruct, indicating that reward-signal densification is complementary to policy-update-level RLVR advances. Code and data are available at https://github.com/tokencraft-lab/SCOPE-RL.
Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Language Models (LLMs) by leveraging direct outcome verification instead of learned reward models. Building on this paradigm, Group Relative Policy Optimization (GRPO) eliminates the need for critic models but suffers from indiscriminate credit assignment for intermediate steps, which limits its ability to identify effective reasoning strategies and incurs overthinking. In this work, we introduce a model-free and verifiable process supervision via probing the model's belief in the correct answer throughout its reasoning trajectory. By segmenting the generation into discrete steps and tracking the conditional probability of the correct answer appended at each segment boundary, we efficiently compute interpretable segment-wise progress measurements to refine GRPO's trajectory-level feedback. This approach enables more targeted and sample-efficient policy updates, while avoiding the need for intermediate supervision derived from costly Monte Carlo rollouts or auxiliary models. Experiments on mathematical and general-domain benchmarks show consistent gains over GRPO across diverse models: up to 2.6-point accuracy improvements and 13.7% reasoning-length reductions on math tasks, and up to 2.4 points and 4% on general-domain tasks, demonstrating strong generalization.
Jingyi Wang, Lei Zhu, Tengjin Weng +8
Tsinghua University · Huawei Noah’s Ark Lab · Shenzhen University
RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling and weighting, leading to inefficient compute allocation. We study GRPO by tracking token log-probabilities, group-normalized advantages, and the induced token-level update weights. This reveals three recurring dynamics as training proceeds: (1) confidence inflation, (2) advantage contraction, and (3) hierarchical convergence. These findings suggest that the utility of each update depends strongly on both question difficulty and the model's current competence. Motivated by this, we propose Confidence and Difficulty-adaptive Policy Optimization (CoDaPO), which assigns each question a bounded value from rollout confidence and empirical difficulty. CoDaPO then uses this value to reweight policy updates and resample high-value learnable questions within mini-batches, thereby increasing discovery within the learnable band under a fixed compute budget. Across twelve benchmarks, CoDaPO consistently improves accuracy over existing RL methods. Our code is publicly available at https://github.com/tmlr-group/CoDaPO.
Zhanke Zhou, Xiangyu Lu, Chentao Cao +4
1TMLR Group, Department of Computer Science, Hong Kong Baptist University · 2Stanford University · 3Sydney AI Centre, The University of Sydney