Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.
Figures & tables
Figure 1: SCAPO achieves the highest AIME accuracy on the two Qwen3 base models. Aggregate AIME 2024–2026 accuracy for SCAPO, GRPO, and FIPO. SCAPO reaches GRPO’s final accuracy in fewer than half as many policy optimization steps and achieves higher final accuracy than baselines at both model scales. Shaded regions mark SCAPO’s semifactual credit augmentation phase.
Figure 2: Semifactual instability varies across tokens. (a) Probability transitions under semifactual prompt interventions. (b) Distribution of mean token drift. (c) Relative category mean drift normalized by the overall mean, with reflection markers and discourse connectives showing higher sensitivity. Results use frozen Qwen3-4B-Base before RL.
Figure 3: Semifactual filtering improves rollout accuracy. Suppressing high-drift token candidates during rollout generation improves Qwen3-4B-Base’s accuracy by 14.2 percentage points on the same 1,000-question diagnostic panel.
Figure 4: Overview of SCAPO. (a) Semifactual prompt interventions are paired with a fixed group of sampled responses. (b) Teacher forcing measures sampled-token probability drift under each intervention. (c) Group-wise normalization and aggregation yield relative stability scores, retaining only their negative part. (d) The scaled stability correction, with gradients stopped, is added to the GRPO advantage to produce token-level advantages for policy optimization.
Benchmark
Base
GRPO
GSPO
SAPO
CF-GRPO
FIPO
SCAPO
Qwen3-4B-Base
In-Distribution
AIME 24–26
7.64
22.08
23.47
24.03
22.64
25.42
27.71
AMC 23–25
37.40
59.74
64.71
65.45
62.06
64.57
65.65
HMMT 25–26
2.08
11.71
13.69
15.18
14.09
15.28
16.17
BRUMO 25
18.54
28.96
33.75
33.54
32.08
35.00
37.71
SMT 25
12.74
27.12
26.06
27.12
26.77
28.18
30.31
Table 1: Overall accuracy on competition-level mathematical reasoning and out-of-distribution benchmarks (%). Base denotes the model before RL. AIME variants aggregate 2024–2026, while HMMT and AMC aggregate 2025–2026 and 2023–2025, respectively. Year-specific results are provided in Appendix F.1 . Bold and underline indicate the best and second-best results, respectively.
Method
AIME 24–26
AMC 23–25
HMMT 25–26
BRUMO 25
SMT 25
Avg.
GRPO
7.71
32.97
1.98
13.54
6.96
12.63
SCAPO
11.88
41.39
4.96
19.17
10.85
17.65
Credit signal source
Random shuffle
8.33
32.78
2.28
15.42
9.43
13.65
Counterfactual
6.32
30.51
1.69
12.92
6.96
11.68
Correction sign
Table 2: Ablations on Qwen3-1.7B-Base. SCAPO uses negative-only corrections.
Figure 5: Training time breakdown of SCAPO on Qwen3-4B-Base. Other includes validation, reward computation, old policy log-probability computation, and miscellaneous overhead.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Classification rule
Token examples
Share (%)
Mean dt
Relative mean
Reflection marker a
Matches the reflection lexicon after lexical normalization.
any , check , Now , However , verify , again
0.29
0.04496
2.74
Discourse connective
Matches the 80-form DiMLex-Eng lexicon after normalization. Excludes reflection markers.
and , for , if , or , Thus , Therefore
2.68
0.03807
2.32
Word
Contains a Unicode letter and matches neither lexical category nor the math-symbol rule.
the , of , is , to , we , in
40.92
0.02506
1.53
Punctuation
All non-whitespace characters have Unicode category P* . Excludes math symbols.
, . : # ), :\n
6.75
0.01956
1.19
Math symbol
Begins with \ or contains only characters from the fixed math/ L a T e X symbol set. Excludes numbers.
\ , = , + , - , { , }
26.62
0.00768
0.47
Number
Matches a signed integer, decimal, or percentage expression.
0 , 1 , 2 , 3 , 4 , 5
15.31
0.00612
0.37
Appendix
Table 4: Token categories and drift statistics underlying Figure 2 (c). Share is the percentage of the 870,586 response token positions assigned to each category. Mean is the category’s average drift dt . Relative mean is this average divided by the overall mean (0.01640). denotes a space and \n a newline.
Figure 6: Token-level semifactual sensitivity in a compound-interest problem.
Figure 7: Token-level semifactual sensitivity in a numeral-base problem.
Metric
Base sampling
d -filtered
Correct answers
158
300
Accuracy (%)
15.80
30.00
Paired correctness outcomes (question counts)
Filtered correct
Filtered incorrect
Baseline correct
98
60
Baseline incorrect
202
640
Appendix
Table 5: Decoding results on 1,000 questions with frozen Qwen3-4B-Base. The lower block reports question counts for the four combinations of baseline and filtered answer correctness.
Hyperparameter
Qwen3-4B-Base
Qwen3-1.7B-Base
Data and rollout settings
Training dataset
DAPO-Math-17K ( Yu et al., 2025 )
Maximum prompt length
1,024
Maximum response length
16,384
Rollout batch size (prompts)
128
Responses per prompt ( G )
8
Appendix
Table 6: Shared and method-specific training hyperparameters. Batch sizes count prompts unless stated otherwise. Each rollout step contains two policy optimization steps. Clipping offsets are relative to 1.
Benchmark
Problems
Samples per problem
AIME 2024 / 2025 / 2026
30 / 30 / 30
16
AMC 2023 / 2024 / 2025
40 / 45 / 42
16
HMMT Feb. 2025 / 2026
30 / 33
16
BRUMO 2025
30
16
SMT 2025
53
16
Omni-Math
2,821
1
Appendix
Table 7: Evaluation benchmarks and sample counts. Slash-separated problem counts follow the listed years. ∗ For GPQA-Diamond, we enumerate all 4!=24 permutations of the multiple-choice options to avoid contamination.
In-Distribution
Out-of-Distribution
Method
AIME
AMC
HMMT
NoOp-AIME
ThinkBench-AIME
24
25
26
23
24
25
25
26
24
25
26
24
25
26
Qwen3-4B-Base
Base
8.33
7.92
6.67
47.66
32.50
32.89
1.04
3.03
8.33
7.29
4.58
8.96
3.75
6.88
GRPO
23.54
23.13
19.58
66.41
55.69
57.74
10.83
12.50
21.25
21.25
19.17
25.42
22.92
19.58
GSPO
27.08
22.92
20.42
73.59
61.39
59.82
10.42
16.67
19.58
21.46
19.38
24.58
22.71
18.96
Appendix
Table 8: Yearly accuracy (%) for Table 1 . Bold and underlining denote the best and second-best results.
Figure 8: Pass@ k on the AIME 2024–2026 and AMC 2023–2025 benchmarks at both model scales.