SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over 100×.
Figures & tables
Gradient Descent with Momentum
TestGRAD
Parameter state. Current parameter vector θt .
Test suite state. Current candidate test suite yt .
Objective. Loss function L(θt) .
Objective. Differential loss L(yt) measuring whether the current test suite has produced a differential outcome.
Retention strength. β∈[0,1] controls how much of the previous velocity is retained.
Support-based retention. The β=MinSupport threshold defines the minimum frequency a sequence must satisfy. Lowering this threshold retains more sequences and gives momentum broader influence.
Momentum. β∇t−1 . The accumulated direction from previous steps.
Momentum analogue. mt=Mineβ(M) . A persistent test update distilled from recurring failure sequences in memory M , after the latest failed update has been appended.
Update with Momentum. ∇t=β∇t−1+(1−β)∂θt∂L . Combines instantaneous gradient with momentum to suppress noisy one-step gradients and reinforce consistent directions.
Gradient analogue. ∇TestGrad=Optimizer(L,yt,mt) . Combines current feedback with momentum to suppress one-off noisy edits and reinforce consistent test-improvement directions.
State update. θt+1=θt−η∇t .
Test update. yt+1=yt+∇TestGrad , where ∇TestGrad is a combination of CRUD operations on the original test suite.
Table 1: Role-level analogy between gradient descent with momentum and test-space optimization. A one-step CRUD edit plays the role of an instantaneous gradient, whereas a high-support recurring failed edit sequence plays the role of momentum carried across attempts.
Figure 1: TestGRAD framework overview. Given an issue and candidate patches from multiple SWE agents, TestGRAD iteratively evolves the repository test suite: the forward pass executes tests to obtain differential loss; failed non-differential updates are stored in failure memory and mined into Failure Pattern Momentum; the backward pass uses the loss and momentum to produce the next CRUD test update.
Figure 2: TestGRAD gradient computation in the backward pass. The derivative step is implemented by an LLM optimizer that converts execution-defined loss and Failure Pattern Momentum into a CRUD update over the repository test suite.
Figure 3: TestGRAD momentum-gradient transition from yt to yt+1 , illustrated on sphinx-doc__sphinx-10673 . (a) The current suite yt checks warning text and yields uniform outcomes ( L(yt)=1 ). (b) Failure memory is mined into mt ; the highlighted statements form the frequent sequence shared by all three failed suites (support 3≥β ), marking this repeated failed direction as a dead end and steering the optimizer toward AST inspection. (c) The updated suite yt+1 probes toctree structure with assert_node() , producing a discriminative signal for the checker. Colors: test momentum .
Base LLM
Augment
Agentless
TRAE
Ours
Pass@1 (%)
Gain (%)
Pass@1 (%)
Gain (%)
Pass@1 (%)
Gain (%)
Pass@1 (%)
Gain (%)
Minimax-m2.7
79.0
- 0.2
80.4
+ 1.2
80.6
+ 1.4
84.2
+ 5.0
GLM-5.1
78.6
- 0.6
80.2
+ 1.0
80.2
+ 1.0
83.6
+ 4.4
DeepSeek-v3.2
79.2
+ 0.0
79.8
+ 0.6
80.6
+ 1.4
84.0
+ 4.8
Table 2: Performance comparison. Gain is computed relative to the baseline Pass@1 (79.2%). Positive gains are shown in green, while negative gains are shown in red.
Figure 4: Pass@1 vs. ensemble size for TestGRAD (green) and TRAE (orange), together with Oracle (blue dashed), Adversary (red dashed), and Expected (purple dotted) bounds. Results shown for Minimax, DeepSeek, and GLM (left to right).
Agent
N
Adv.
Exp.
Oracle
P@1
Livesweagent
1
79.2
79.2
79.2
79.2
+Sonar
2
76.8
79.8
82.8
81.2
+OpenHands
3
74.6
79.3
83.4
81.2
+Trae_doubao
4
70.2
79.2
87.0
84.2
+Gemini
5
68.4
79.1
88.0
83.8
+Atlassian
6
66.8
78.7
88.0
83.4
Table 3: Pass@1 across ensemble sizes ( TestGRAD , Minimax-m2.7).
Model
Setting
P@1
Δ
TestGRAD
Complete
84.2
–
\textscTestGRADw/o Read
w/o suite understanding
81.2
-3.0
\textscTestGRADw/o Update
w/o updating tests
81.8
-2.4
\textscTestGRADw/o R+U
w/o Read and Update
79.4
-4.8
\textscTestGRADw/o Diff Loss
pass-count selection
80.8
-3.4
\textscTestGRADw/o Momentum
w/o failure momentum
82.6
-1.6
Table 4: Ablation study on SWE-bench Verified (Minimax-m2.7).
Figure 5: Average token consumption of Raw Memory (dark gray) vs. Momentum (coral) across optimization iterations, shown on a log scale for DeepSeek, GLM, and Minimax. Italic labels denote representative compression ratios (Raw / Momentum); ratios above 100× are capped as >100x .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: (A) CD-only generation produces a reproduction script that must hand-roll all setup boilerplate and tests the old exception-raising behavior, making it non-discriminative; tests are deleted by name heuristic, discarding valid ones. (B) TestGRAD applies a CRUD gradient: it reads the test suite first to extract infrastructure ( @isolate_apps , SimpleTestCase , .check() convention), creates a grounded test reusing that infrastructure ( blue = reused from Read), updates the expected model state in test_render , and semantically deletes test_missing_parent_link whose expected exception is no longer raised after the fix. Instance: django__django-12325 .
Table 1: Frequency of the six retained test-suite understanding aspects across 200 sampled SWE-bench instances (SWE-bench Verified excluded). Frequency is the fraction of instances where the aspect is present in the target test file. The six discarded aspects each fell below 50%.
Metric
Method
Minimax-m2.7
DeepSeek-v3.2
GLM-5.1
Avg. Time (min)
Augment
1.3
1.8
1.4
Agentless
7.2
8.6
7.5
TRAE
19.3
26.7
20.8
TestGRAD
12.2
23.0
20.7
Avg. Tokens (K)
Augment
2.8
3.2
2.9
Agentless
14.8
16.2
15.3
Appendix
Table 2: Average per-instance overhead across methods and base LLMs. Time is in minutes; tokens are in thousands (K). TestGRAD values are measured from run logs; baseline values are estimated from conversation logs and API call traces.