SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over 100×.
Figures & tables
Gradient Descent with Momentum
TestGRAD
Parameter state. Current parameter vector θt .
Test suite state. Current candidate test suite yt .
Objective. Loss function L(θt) .
Objective. Differential loss L(yt) measuring whether the current test suite has produced a differential outcome.
Retention strength. β∈[0,1] controls how much of the previous velocity is retained.
Support-based retention. The β=MinSupport threshold defines the minimum frequency a sequence must satisfy. Lowering this threshold retains more sequences and gives momentum broader influence.
Momentum. β∇t−1 . The accumulated direction from previous steps.
Momentum analogue. mt=Mineβ(M) . A persistent test update distilled from recurring failure sequences in memory M , after the latest failed update has been appended.
Update with Momentum. ∇t=β∇t−1+(1−β)∂θt∂L . Combines instantaneous gradient with momentum to suppress noisy one-step gradients and reinforce consistent directions.
Gradient analogue. ∇TestGrad=Optimizer(L,yt,mt) . Combines current feedback with momentum to suppress one-off noisy edits and reinforce consistent test-improvement directions.
State update. θt+1=θt−η∇t .
Test update. yt+1=yt+∇TestGrad , where ∇TestGrad is a combination of CRUD operations on the original test suite.
Table 1: Role-level analogy between gradient descent with momentum and test-space optimization. A one-step CRUD edit plays the role of an instantaneous gradient, whereas a high-support recurring failed edit sequence plays the role of momentum carried across attempts.
Figure 1: TestGRAD framework overview. Given an issue and candidate patches from multiple SWE agents, TestGRAD iteratively evolves the repository test suite: the forward pass executes tests to obtain differential loss; failed non-differential updates are stored in failure memory and mined into Failure Pattern Momentum; the backward pass uses the loss and momentum to produce the next CRUD test update.
Figure 2: TestGRAD gradient computation in the backward pass. The derivative step is implemented by an LLM optimizer that converts execution-defined loss and Failure Pattern Momentum into a CRUD update over the repository test suite.
Figure 3: TestGRAD momentum-gradient transition from yt to yt+1 , illustrated on sphinx-doc__sphinx-10673 . (a) The current suite yt checks warning text and yields uniform outcomes ( L(yt)=1 ). (b) Failure memory is mined into mt ; the highlighted statements form the frequent sequence shared by all three failed suites (support 3≥β ), marking this repeated failed direction as a dead end and steering the optimizer toward AST inspection. (c) The updated suite yt+1 probes toctree structure with assert_node() , producing a discriminative signal for the checker. Colors: test momentum .
Base LLM
Augment
Agentless
TRAE
Ours
Pass@1 (%)
Gain (%)
Pass@1 (%)
Gain (%)
Pass@1 (%)
Gain (%)
Pass@1 (%)
Gain (%)
Minimax-m2.7
79.0
- 0.2
80.4
+ 1.2
80.6
+ 1.4
84.2
+ 5.0
GLM-5.1
78.6
- 0.6
80.2
+ 1.0
80.2
+ 1.0
83.6
+ 4.4
DeepSeek-v3.2
79.2
+ 0.0
79.8
+ 0.6
80.6
+ 1.4
84.0
+ 4.8
Table 2: Performance comparison. Gain is computed relative to the baseline Pass@1 (79.2%). Positive gains are shown in green, while negative gains are shown in red.
Figure 4: Pass@1 vs. ensemble size for TestGRAD (green) and TRAE (orange), together with Oracle (blue dashed), Adversary (red dashed), and Expected (purple dotted) bounds. Results shown for Minimax, DeepSeek, and GLM (left to right).
Agent
N
Adv.
Exp.
Oracle
P@1
Livesweagent
1
79.2
79.2
79.2
79.2
+Sonar
2
76.8
79.8
82.8
81.2
+OpenHands
3
74.6
79.3
83.4
81.2
+Trae_doubao
4
70.2
79.2
87.0
84.2
+Gemini
5
68.4
79.1
88.0
83.8
+Atlassian
6
66.8
78.7
88.0
83.4
Table 3: Pass@1 across ensemble sizes ( TestGRAD , Minimax-m2.7).
Model
Setting
P@1
Δ
TestGRAD
Complete
84.2
–
\textscTestGRADw/o Read
w/o suite understanding
81.2
-3.0
\textscTestGRADw/o Update
w/o updating tests
81.8
-2.4
\textscTestGRADw/o R+U
w/o Read and Update
79.4
-4.8
\textscTestGRADw/o Diff Loss
pass-count selection
80.8
-3.4
\textscTestGRADw/o Momentum
w/o failure momentum
82.6
-1.6
Table 4: Ablation study on SWE-bench Verified (Minimax-m2.7).
Figure 5: Average token consumption of Raw Memory (dark gray) vs. Momentum (coral) across optimization iterations, shown on a log scale for DeepSeek, GLM, and Minimax. Italic labels denote representative compression ratios (Raw / Momentum); ratios above 100× are capped as >100x .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: (A) CD-only generation produces a reproduction script that must hand-roll all setup boilerplate and tests the old exception-raising behavior, making it non-discriminative; tests are deleted by name heuristic, discarding valid ones. (B) TestGRAD applies a CRUD gradient: it reads the test suite first to extract infrastructure ( @isolate_apps , SimpleTestCase , .check() convention), creates a grounded test reusing that infrastructure ( blue = reused from Read), updates the expected model state in test_render , and semantically deletes test_missing_parent_link whose expected exception is no longer raised after the fix. Instance: django__django-12325 .
Table 1: Frequency of the six retained test-suite understanding aspects across 200 sampled SWE-bench instances (SWE-bench Verified excluded). Frequency is the fraction of instances where the aspect is present in the target test file. The six discarded aspects each fell below 50%.
Metric
Method
Minimax-m2.7
DeepSeek-v3.2
GLM-5.1
Avg. Time (min)
Augment
1.3
1.8
1.4
Agentless
7.2
8.6
7.5
TRAE
19.3
26.7
20.8
TestGRAD
12.2
23.0
20.7
Avg. Tokens (K)
Augment
2.8
3.2
2.9
Agentless
14.8
16.2
15.3
Appendix
Table 2: Average per-instance overhead across methods and base LLMs. Time is in minutes; tokens are in thousands (K). TestGRAD values are measured from run logs; baseline values are estimated from conversation logs and API call traces.
Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt's verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
Evaluating software engineering capabilities has become a core component of modern large language models (LLMs); however, the key bottleneck hindering further scaling lies not in the scarcity of high-quality solutions, but in the lack of high-quality test suites. Test suites are indispensable both for synthesizing program repair trajectories and for providing precise feedback signals in reinforcement learning. Unfortunately, due to the high cost and difficulty of annotation, high-quality test suites have long been hard to obtain, while those automatically generated by LLMs tend to be superficial and lack sufficient discriminative power. As a first step toward constructing high-quality test suites, we introduce SWE-Mutation, a benchmark for evaluating LLM-generated test suites. The benchmark characterizes test suites by introducing systematically mutated solutions that attempt to ``fool'' the test suites and pass validation. We further propose an agentic, language-agnostic framework for automatically generating complex mutants. Our benchmark consists of 2,636 mutated variants derived from 800 original instances and includes a multilingual subset spanning nine programming languages. Experiments on seven LLMs reveal that even DeepSeek-V3.1 achieves only 10.20% verification and 36.15% detection rates, highlighting the inadequacy of current LLMs. Additionally, our agentic mutation strategy enhances realism, reducing average detection rates from 71.04% to 39.81% compared to conventional methods. These findings expose persistent deficiencies in the ability of current LLMs to generate reliable and discriminative test suites.
Yuxuan Sun, Yuze Zhao, Yufeng Wang +6
1State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China · 2Independent Researcher. · 3Beihang University. +3
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion. Early termination, however, risks interrupting trajectories that would otherwise succeed; conversely, an unsuccessful trajectory may still contain useful repository edits. We present FailFast-RestartSmart, a two-stage controller for a single active trajectory. FailFast is a lightweight 0.6B monitor trained with terminal and dense fail-to-pass supervision to predict failure from observable prefixes without policy logits or hidden states. Upon an alarm, RestartSmart launches a fresh same-policy rollout without prior prompt history and offers the interrupted repository diff as an optional overlay that the agent may inspect, apply, or discard. On SWE-bench Verified, a monitor trained solely on Qwen3.6-27B trajectories transfers to three other policies, including a closed-API model, and saves 14.6%-20.4% of execution tokens at a target 5% false-positive rate; on Qwen3.6-27B, its 20.4% saving exceeds the 12.5% achieved by our per-step AgentStop adaptation. At a target 25% false-positive rate, RestartSmart raises Qwen3.6-27B resolution from 66.6% to 71.8%, whereas cold restart reaches only 66.8%. Together, these results support early termination with sequential same-policy recovery.
Chenyu Wang, Yunbo Lyu, Junda He +4
Singapore Management University · University of Alberta · Nanjing University of Science and Technology +1