Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt's verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
Figures & tables
Figure 1: Comparison of trajectory execution in GRPO and FC-SWE . During training, our GRPO baseline samples independent attempts, whereas FC-SWE launches a recovery trajectory after failure, conditioned on the failed patch and verifier output. Each executed trajectory retains its own verifier reward.
Figure 2: End-to-end FC-SWE training system. Rollout workers execute verifier-delimited attempts; failed chains are reset and resumed with bounded failure evidence. Executed trajectories are packed with inactive transport padding, assigned trajectory-local rewards and active-set advantages, and optimized with a token-normalized policy update.
Model / policy
Resolved@1
Resolved@2
Conditional recovery
Qwen3.5-4B ( Qwen Team, 2026a )
27.9%
36.3%
11.6%
Qwen3.5-9B ( Qwen Team, 2026b )
41.2%
50.4%
15.6%
Nemotron-3-Nano-4B ( NVIDIA, 2026 )
28.2%
35.8%
10.6%
Qwen3.5-4B RL training
GRPO ( Shao et al., 2024 )
38.9%
48.5%
15.7%
FC-SWE (ours)
41.7%
52.8%
19.0%
Table 1: Results on all 500 SWE-bench Verified tasks. All results are evaluated by us. Resolved@1 and Resolved@2 are first-attempt and cumulative two-attempt success, averaged over three decoding seeds (avg@3). Conditional recovery pools second-attempt successes over initial failures across seeds. Retries use failure feedback. Higher is better.
Figure 3: Online training success during optimization. (a) First-attempt and cumulative two-attempt success of FC-SWE , compared with single-attempt GRPO; shading shows the gain from recovery. (b) Ablations of failure context and reward design in FC-SWE . Curves use five-step moving averages, except for the collapse segment.
Training variant
Resolved@1
Resolved@2
Conditional recovery
A. Training-time failure context
FC-SWE
41.7%
52.8%
19.0%
w/o failure context
40.4%
47.9%
12.6%
B. Reward assignment and runtime penalties
FC-SWE
43.2%
54.8%
20.4%
w/o trajectory-local rewards
41.4%
51.8%
17.7%
Table 2: Training ablations with failure context retained at evaluation. Block A uses three decoding seeds; Block B uses one. Metric definitions and aggregation follow Table 1 .
Figure 4: Failure-evidence ablations. (a) Feedback retries versus resampling at a matched two-attempt budget; labels show differences in percentage points. (b) Second-attempt context ablations using the final FC-SWE policy.
Figure 5: Test-time scaling and token efficiency (three-seed means). (a) Resolved@ K versus attempt budget; the dashed line marks the two-attempt training horizon. (b) Resolved@ K versus estimated cumulative tokens per task. Token accounting is detailed in Appendix E .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Value
Policy initialization
Qwen3.5-4B
Training / evaluation harness
OpenHands CodeActAgent / SWE-agent
Context budget
65,536 tokens
Maximum model–tool interactions
100 per attempt
Training attempt budget
R=2
Tasks / initial chains per task
16 / 8
Appendix
Table 3: Main training and evaluation configuration.
Figure 6: Selected SWE-bench Verified cases comparing GRPO and FC-SWE . (a) Second-attempt responses to each policy’s own failed patch and verifier feedback. (b) First-attempt patches that both pass the target test but differ in regression-test failures.
Checkpoint
Resolved@1
Recovered
Recovery
Resolved@2
Shared, step 40
37.6%
41
13.1%
45.8%
Shared, step 45
39.4%
41
13.5%
47.6%
Shared, step 55
42.2%
47
16.3%
51.6%
Shared, step 70
41.4%
52
17.7%
51.8%
Local, step 50
42.4%
53
18.4%
53.0%
Local, step 67
43.2%
58
20.4%
54.8%
Appendix
Table 4: Checkpoint evaluations on all 500 tasks using one decoding seed. Recovered counts tasks first resolved on attempt 2.
K
Base
GRPO
FC-SWE
1
150/124/145 (27.9%)
206/179/198 (38.9%)
216/207/203 (41.7%)
2
183/168/193 (36.3%)
257/228/242 (48.5%)
274/261/257 (52.8%)
3
221/193/229 (42.9%)
278/255/278 (54.1%)
289/280/288 (57.1%)
4
241/227/253 (48.1%)
292/269/297 (57.2%)
305/296/309 (60.7%)
5
251/243/266 (50.7%)
304/284/313 (60.1%)
318/306/321 (63.0%)
6
260/258/277 (53.0%)
310/298/317 (61.7%)
329/320/330 (65.3%)
Appendix
Table 5: Cumulative solved tasks out of 500. Each cell gives counts for decoding seeds 2/3/4, followed by the mean resolution rate.
Figure 7: Additional test-time scaling views. (a) Initial failures recovered within K attempts, pooled across seeds. (b) Cumulative resolution versus estimated token cost, reproduced for reference; markers denote evaluated attempt budgets.
K
Tokens/task (M)
Attempts/task
Synth. hours
Resolved@ K
Base
1
2.36
1.00
33
27.9%
2
4.06
1.72
57
36.3%
3
5.59
2.36
78
42.9%
4
6.98
2.93
96
48.1%
5
8.22
3.45
113
50.7%
Appendix
Table 6: Per-budget inference accounting, averaged over three decoding seeds. Token costs are estimated; the hours column is a synthesized estimate for all 500 tasks, not measured latency.
Figure 8: Training execution statistics per active attempt: (a) tool-call turns, (b) generated tokens, and (c) logged truncation. Light lines show individual steps; dark lines show centered five-step averages with shortened endpoint windows.
State Key Laboratory of Networking and Switching Technology Beijing University of Posts and Telecommunications, Beijing 100876, China · University of Luxembourg, Luxembourg