Organizations: State Key Laboratory of Networking and Switching Technology Beijing University of Posts and Telecommunications, Beijing 100876, China · University of Luxembourg, Luxembourg
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.
Figures & tables
GRPO
SWE-Trace
CRR
Device-time (H100-hours)
188
233
321
Node wall-clock (hours)
23.5
29.1
40.1
Verified pass@1 (%)
36.4±0.8
39.1±0.7
41.7±0.6
Table 1: Full-budget resource totals for the three central training comparisons. Device-time includes on-policy rollout, counterfactual replay or PRM scoring, and learner updates. These settings do not have equal total compute. Wall-clock values are rounded to one decimal place.
Method
SWE-bench Verified
SWE-bench Live
SWE-rebench
Δ vs. GRPO
SWE-agent (zero-shot)
22.6±1.4
18.9±1.5
16.4±1.6
−13.8
OpenHands (zero-shot)
23.4±1.4
19.7±1.5
17.0±1.6
−13.0
SWE-Gym SFT
28.1±1.2
24.0±1.3
21.5±1.4
−8.3
SWE-Fixer
30.7±1.1
25.8±1.3
22.8±1.4
−5.7
RLEF (outcome-only PPO)
33.5±1.0
28.6±1.2
25.4±1.3
−2.9
Murphy
35.2±0.9
30.0±1.1
26.7±1.2
−1.2
Table 2: Pass@1 (%) with the same 14B backbone and scaffold. Values are means ± seed standard deviations over three seeds. Bold denotes the best non-stacked result; underlining denotes the best result within this table. The final column is the Verified difference from vanilla GRPO. These endpoint runs are not uniformly compute-matched.
Method
Fork use
Pass@1 (%)
Vanilla GRPO
None
33.3±0.7
Search-Select-GRPO
Sequence selection
34.4±0.8
CRR
Credit assignment
36.2±0.6
Table 3: Two uses of matched fork executions. Every arm receives 48 H100-hours per seed; greedy evaluation covers all 500 Verified instances. Values are means ± seed standard deviations over three seeds.
Configuration
none
TSR
PRM
TSR+PRM
Base GRPO
36.4
38.0
39.1
40.0
Base GRPO + CRR
41.7
43.6
44.9
45.6
Table 4: Composition on Verified, measured by pass@1 (%). TSR, PRM, and their combination are added to the base GRPO configuration. The 44.9% result uses CRR+PRM; 45.6% uses CRR+TSR+PRM.
Method
Greedy pass@1
Sampled pass@1
pass@8
all-8
Vanilla GRPO
36.4
35.9
53.8
18.7
SWE-Trace
39.1
38.6
55.6
21.2
CRR
41.7
41.2
57.7
23.9
Table 5: Sampling evaluation on all 500 Verified instances using eight completions per instance ( T=0.6 ). Reported percentages are averaged over three full-budget training seeds. Greedy pass@1 is included for context.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Step
200
1,000
2,000
3,000
4,000
Cov^ (mean)
0.044
0.071
0.103
0.119
0.128
Cov^ (5th pct.)
0.018
0.029
0.046
0.059
0.068
Cov^ (95th pct.)
0.094
0.131
0.171
0.183
0.198
Appendix
Table 6: Empirical task/prefix-bucket covariance Cov^(R(τ),R(τ~t)∣B) at five checkpoints during training. The covariance remains strictly positive throughout, with the average value increasing as the policy becomes more confident in its prefix decisions.
Table 7: Default hyperparameters for the full-budget CRR runs. Ablations vary their stated factor, and equal-device-time controls may complete fewer trajectories or updates.
Repository
Tasks evaluated
Discard rate (%)
sympy
482
4.1
django
591
11.7
scikit-learn
277
5.4
pytest
162
7.4
matplotlib
191
6.8
sphinx
219
9.6
Appendix
Table 8: Discard rate of the three-rerun determinism filter for the largest repositories in the upstream SWE-Gym and SWE-rebench candidate pools. Repositories with networked test fixtures dominate the discards.
Concurrency W
1
2
4
8
16
Relative fork-pool throughput
1.00
1.75
2.89
4.25
4.46
Verified pass@1 (%)
41.6
41.7
41.7
41.7
41.6
Appendix
Table 9: Fork-pool profiling throughput and pass-rate as a function of counterfactual-worker concurrency W . Throughput is normalised to W=1 inside the profiling harness. These profiling results are not an end-to-end device-time ledger or evidence of invariance under arbitrary changes in concurrency.
Method
Wrong loc.
Partial fix
Regression
Harness subv.
Other
Vanilla GRPO
34.2
39.7
19.4
6.7
0.0
SWE-Trace
33.6
35.1
24.5
6.8
0.0
CRR (ours)
39.5
24.2
11.3
6.6
18.4
CRR + SWE-Trace
40.8
21.9
9.1
6.8
21.4
Appendix
Table 10: Error-mode distribution among failed instances on SWE-bench Verified , normalised to percentages of all failures of the corresponding method. The table reports conditional shares among failures; Table 11 converts the main comparison to approximate absolute rates over all evaluated instances.
Method
Wrong loc.
Partial fix
Regression
Harness subv.
Other
Vanilla GRPO
21.75
25.25
12.34
4.26
0.00
CRR (ours)
23.03
14.11
6.59
3.85
10.73
Appendix
Table 11: Approximate absolute error-mode rates over all SWE-bench Verified instances, obtained by multiplying Table 10 ’ conditional failure shares by the corresponding failure rates from Table 2 .
Component
Vanilla GRPO
CRR (ours)
Δ
Forward rollout (32 turns)
78.0
78.0
+0.0
Test execution
9.4
9.4
+0.0
Counterfactual fork ( k=4 )
0.0
36.2
+36.2
Counterfactual rollout
0.0
28.1
+28.1
Counterfactual test execution
0.0
6.3
+6.3
Policy update (GRPO step)
12.7
12.9
+0.2
Appendix
Table 12: Critical-path wall-clock latency per 64-trajectory rollout-dispatch microbatch on 8× H100 (seconds). The reported CRR microbatch latency is 1.71× that of vanilla GRPO. This local profile does not replace the end-to-end device-time accounting in Table 14 or the training-window comparison in Table 15 .
Quantity
Main value
Meaning
Training corpus
3,638 tasks
post-filter admitted deterministic tasks
SWE-rebench split
disjoint
train/eval separated before filtering and decontaminated
On-policy trajectories
24,000
newly collected policy rollouts per full-budget CRR run/seed
Fork budget
k=4
maximum selected decision points per trajectory
Counterfactual samples
K=1
main setting; K=2,4 only in ablations
Main counterfactual tuples
about 96,000
nominal maximum 24,000×k×K per seed
Appendix
Table 13: Configuration and accounting conventions for the full-budget CRR run. Method-specific device-time components are reported in Table 14 ; equal-wall-clock checkpoints are reported in Table 15 .
Accounting item
Vanilla GRPO
SWE-TRACE
CRR
On-policy rollout (H100-hours)
164
164
164
Counterfactual fork pool (H100-hours)
0
0
133
PRM scoring / auxiliary work (H100-hours)
0
45
0
Learner updates (H100-hours)
24
24
24
Total device-time (H100-hours)
188
233
321
Wall-clock on one 8× H100 node (hours)
23.5
29.1
40.1
Appendix
Table 14: Full-budget resource accounting and SWE-bench Verified pass@1. Performance is the mean ± seed standard deviation over three training seeds. Device-time includes all listed worker roles. Wall-clock durations in this table are rounded to one decimal place; the exact CRR endpoint used in the matched-time comparison is 40.125 hours.
Fraction of CRR training window
25%
50%
75%
100%
Wall-clock (hours)
10.031
20.063
30.094
40.125
GRPO pass@1 (%)
34.6
36.1
36.5
36.7
GRPO completed trajectories
11,648
20,992
29,632
37,824
GRPO optimizer updates
1,941
3,499
4,939
6,304
GRPO generated tokens (M)
28.70
57.40
86.11
114.76
GRPO effective tokens (M)
23.54
47.07
70.61
94.10
Appendix
Table 15: Full-budget matched-wall-clock comparison on the same 8× H100 node. Pass@1 is the three-seed mean on SWE-bench Verified ; trajectory, update, and token rows give the cumulative ledger at the stated checkpoints. Token counts are in millions. All four checkpoints are observed.
Fork-cost multiplier
m=1
m=2
m=4
CRR pass@1 (%)
36.2
34.8
33.2
Vanilla GRPO pass@1 (%)
33.4
33.4
33.4
CRR minus GRPO (percentage points)
+2.8
+1.4
−0.2
CRR completed on-policy trajectories
≈6,000
4,312
≈2,760
Appendix
Table 16: Fork-overhead sensitivity under the reduced-budget protocol: 48 H100-hours per arm, one training seed, and greedy pass@1 on all 500 Verified instances. The GRPO reference is the same 33.4% run for each cost multiplier.
Method
Use of forks
Pass@1 (%)
Vanilla GRPO
None
33.3±0.7
Search-Select-GRPO
Select the training sequence
34.4±0.8
CRR
Credit assignment on realized actions
36.2±0.6
Appendix
Table 17: Matched Search-Select comparison. All arms use 48 H100-hours per seed and greedy evaluation on all 500 Verified instances. Values are the mean ± seed standard deviation over three training seeds. CRR and Search-Select use the same fork count per realized trajectory; GRPO has no forks.
Method
Pass@1 greedy
Pass@1 sampled
Pass@2
Pass@4
Pass@8
All-8
Vanilla GRPO
36.4
35.9
43.4
48.8
53.8
18.7
SWE-TRACE
39.1
38.6
45.7
50.8
55.6
21.2
CRR
41.7
41.2
48.1
53.1
57.7
23.9
Appendix
Table 18: Multiple-sample evaluation on SWE-bench Verified using full-budget final checkpoints and three training seeds. Each instance receives eight independent samples at temperature 0.6. All values are percentages; the greedy column is evaluated separately from the sampled columns.
Configuration
Forks per trajectory
Pass@1 (%)
Relative advantage SNR
K=1,k=4 (default)
4
36.2
2.0×
K=2,k=4 (twice as many forks)
8
36.6
2.5×
K=2,k=2 (matched forks)
4
35.8
2.3×
K=4,k=1 (matched forks)
4
35.0
2.6×
Appendix
Table 19: Counterfactual-sample sensitivity under the reduced-budget protocol: 48 H100-hours per arm, one training seed, and greedy evaluation on all 500 Verified instances. Advantage SNR uses the main evaluation’s diagnostic and is expressed relative to GRPO.
Injected nondeterminism ρ
0%
5%
10%
CRR, K=1 : pass@1 (%)
36.2
35.6
34.8
CRR, K=2 : pass@1 (%)
—
36.0
35.4
Vanilla GRPO reference: pass@1 (%)
33.4
33.4
33.4
Mean advantage shift
0
0.01
0.02
Appendix
Table 20: Controlled replay-noise sensitivity under the reduced-budget protocol: 48 H100-hours per arm, one training seed, and greedy pass@1 on all 500 Verified instances. The mean-advantage-shift row is a diagnostic of the injected perturbation, not a pass-rate. A dash denotes an unreported configuration.
Alternative class
Fork frequency (%)
Mean contrast
Plug-in share (%)
Immediately invalid
12
+0.15
8
Valid, branch fails
46
+0.40
83
Valid, branch succeeds
42
−0.05
9
Appendix
Table 21: Composition of 20,000 logged counterfactual forks. This is a log diagnostic, not a separate 500-instance evaluation or a multi-seed performance comparison. The final column is the normalized class-mean plug-in quantity pc∣Δˉc∣/∑jpj∣Δˉj∣ , expressed as a percentage.
Proposal
Pass@1 (%)
Unfiltered counterfactual proposal
36.2
Validity-filtered counterfactual proposal
36.4
Appendix
Table 22: Validity-filtered proposal under the reduced-budget protocol: 48 H100-hours per arm, one training seed, and greedy evaluation on all 500 Verified instances.
Selector
Hits gold-patch file on first touch (%)
Selected steps in top-decile ∣A^∣ (%)
Mean ∣A^∣ at selected steps
Random- k
13
10
0.11
Stride- k
15
11
0.11
Entropy only
30
24
0.18
Entropy + tool type
41
28
0.21
Appendix
Table 23: Selector alignment with decision-relevance proxies. The first numeric column uses 2,000 logged training trajectories; the remaining columns use the shared 300-trajectory probe. These are diagnostics of training trajectories, not pass@1 measurements on the 500-instance test set.