Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6--81.0% and total verification tokens, which include test generation, by 49.1--62.1% at K=16, while hybrid Best@16 matches or improves on each agent's reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.
Figures & tables
Figure 1: Existing hybrid verification and the workflow with LatentSift. The existing pipeline follows the hybrid verification design used by R2E-Gym and DeepSWE-Preview ( Jain et al., 2025 ; Luo et al., 2025 ) . Both pipelines select a patch from the same pool of K generated candidates. The red and green dashed boxes mark inference-time verification for the existing pipeline and ours, respectively. The existing pipeline scores all candidates with an LLM-based execution-free (EF) verifier in Stage 1 and reuses these scores in Stage 4. LatentSift replaces Stage 1 with a filter that reuses cached policy states without additional LLM tokens. Both pipelines retain half the candidates before regression testing and LLM-generated testing; ours applies the final EF verifier only to the remaining candidates. Token counts show average costs per task instance on DeepSWE-Preview at K=16 , with candidate generation reported separately.
Figure 2: LatentSift verification design. Left: Reasoning, observation, and function states from outcome-labeled training trajectories form positive and negative banks. Inference reuses candidate states from ordinary policy forward passes, and Stage 1 averages the distance score with the linear scorer’s within-instance rank to retain the top half of candidates. Lower right: Distance scoring: (1) retrieve the nearest positive and negative states and compute mtc=dnegc−dposc . (2) aggregate by step-weighted mean; (3) rank within each task instance and channel; (4) take the minimum channel rank. Upper right: Average verification tokens per task instance for the original and token-free cascades at K=16 (Section 4.2.1 ).
DeepSWE Agent
R2EGym Agent
CWM Agent
Stage 1 verifier
Best@16
Token cost (K)
Best@16
Token cost (K)
Best@16
Token cost (K)
EF
Total
EF
Total
EF
Total
DeepSWE-Verifier
59.26
826.7
1035.1
49.70
367.4
599.4
54.21
692.6
930.0
R2EGym-Verifier
57.95
774.1
982.0
46.56
366.2
599.6
53.51
707.1
944.6
DEV matching-pairs
58.75
826.7
1037.6
48.38
367.4
599.4
53.11
692.6
925.6
AgentPRM
59.76
810.9
1022.8
46.56
358.3
586.8
53.31
677.3
909.8
Table 1: Cross-agent hybrid results on SWE-bench Verified at K=16 . Best@16 is the resolution rate (%). Token cost gives EF-verifier (EF) and total verification tokens per task instance, in thousands.
Stage-1 filter
K=4
K=6
K=8
K=10
K=12
K=14
K=16
LLM EF verifier
97.67
94.67
95.79
96.37
96.81
97.13
97.44
LatentSift (ours)
97.07
93.77
94.95
95.97
96.74
97.28
97.72
Random filtering
92.31
85.91
88.27
90.03
91.49
92.50
93.19
Table 2: Oracle retention (%) after Stage 1 on DeepSWE-Preview at even K , retaining max(3,K/2) candidates, over the same 200 draws as Figure 4 .
Distance channels
R+O+F
R+O
R+F
O+F
R
O
F
DeepSWE Agent
60.06
59.26
59.66
59.86
59.46
59.26
59.46
R2EGym Agent
47.73
46.60
46.90
47.10
46.71
46.46
47.25
Table 3: Channel selection ablation of LatentSift at K=16 : hybrid Best@16 (%) with the distance branch restricted to subsets of the reasoning (R), observation (O), and function (F) channels.
DeepSWE Agent
R2EGym Agent
Distance score
Best@16
Ret.
Best@16
Ret.
dnegc−dposc
60.06
97.72
47.73
92.13
dposc
58.05
97.30
44.16
89.31
dnegc
59.86
97.44
47.30
91.48
Table 4: Distance construction ablation at K=16 (%): hybrid Best@16 and oracle retention (Ret.).
DeepSWE Agent
R2EGym Agent
Stage-1 rule
Best@16
Ret.
Best@16
Ret.
LatentSift (step-weighted mean)
60.06
97.72
47.73
92.13
Uniform step mean
59.66
97.44
46.36
90.16
Trajectory length, tokens
57.24
94.87
45.65
91.15
Trajectory length, steps
58.39
95.29
47.60
92.40
Table 5: Naive Stage-1 filtering rules against LatentSift at K=16 (%): hybrid Best@16 and oracle retention (Ret.).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
DeepSWE Agent
R2EGym Agent
CWM Agent
Stage 1 verifier
K=4
K=8
K=12
K=16
K=4
K=8
K=12
K=16
K=4
K=8
K=12
K=16
DeepSWE-Verifier
53.99
57.10
58.43
59.26
42.65
46.11
48.23
49.70
52.13
52.93
53.80
54.21
R2EGym-Verifier
53.67
56.38
57.59
57.95
42.35
45.04
46.27
46.56
52.20
53.28
53.77
53.51
DEV matching-pairs
53.71
56.70
57.82
58.75
42.49
45.58
47.38
48.38
52.13
52.66
53.07
53.11
AgentPRM
53.55
56.70
58.40
59.76
41.47
44.22
45.74
46.56
51.96
52.58
52.85
53.31
LatentSift (ours)
53.45
56.92
58.64
60.06
41.79
44.23
46.12
47.73
51.75
52.65
53.60
54.21
Appendix
Table 6: Hybrid Best@ K (%) of the cross-agent comparison at K=4 , 8, 12, and 16. The K=16 column repeats Table 1 .
R2EGym-32B
R2EGym-14B
Stage 1 verifier
Best@16
Token cost (K)
Best@16
Token cost (K)
EF
Total
EF
Total
R2EGym-Verifier
46.56
366.2
599.6
43.28
344.2
575.3
LatentSift (ours)
47.73
76.0
304.9
42.95
65.5
291.2
Appendix
Table 7: Cross-policy results on R2EGym Agent at K=16 . Best@16 is the resolution rate (%). Token cost gives EF-verifier (EF) and total verification tokens per task instance, in thousands. Both cascades use R2EGym-Verifier in Stage 4.