Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifies the first complete solution in each trajectory, determines whether it is objectively correct, and uses this ground truth to quantify the agent's subsequent verification and recovery behavior. Applying it to ten terminal agents on TerminalBench2.1, we find that verification is nearly universal after a complete candidate is formed, yet only 61.43% of incorrect candidates are detected and only 49.36% of detected errors are successfully repaired. These results show that the main weakness in self-verification lies not in initiating verification, but in detecting and repairing errors. Motivated by these findings, we propose Student-Conditioned Verification Distillation (SCVD), which lets the student first produce a candidate solution and distills a stronger teacher's subsequent verification and recovery from the same interaction context. Across three Qwen3.5 backbones, SCVD improves \textsc{Pass@1} on TerminalBench2.1 by 9.74--16.85 percentage points over the corresponding base models and by 4.49--8.61 points over the standard full-trajectory distillation, while avoiding the pronounced out-of-distribution degradation of full-trajectory distillation on SWE-bench Verified.
Figures & tables
Figure 1: Overview of solution-level self-verification and SCVD. (a) After forming a candidate solution, the agent performs task-relevant checks, repairs detected errors, and rechecks the result before termination. (b) SCVD uses a teacher to continue from the student-generated candidate state to verify and revise the solution. After filtering successful trajectories and removing the temporary verification scaffold, the SFT loss is applied only to the teacher continuation.
Candidate correctness
z=error
z=no-error
z=unresolved
Incorrect ( yc=0 )
Detected error
Missed error
No usable outcome
Correct ( yc=1 )
False alarm
Correct pass
No usable outcome
Table 1: Verification outcomes and diagnostic metrics for solution-level self-verification. Here, yc and yf denote the correctness of the candidate and final solutions, respectively; z denotes the verification outcome; and Ic and Iv indicate the presence of a complete candidate and whether verification is attempted.
Figure 2: Self-verification diagnostics across ten terminal agents on TerminalBench2.1. Metric definitions are given in Table 1 . (a) Diagnostic results across models; gray bars show the range across models, vertical ticks mark the mean for each metric, and the rightmost numbers report corresponding mean values. (b) Improvement from initial candidates to final solutions. (c) Relationship between repair success rate and final task accuracy across models.
TerminalBench2.1
SWE-bench Verified (OOD)
Model
Pass@1 ↑
Pass@3 ↑
Δ
Pass@1 ↑
Δ
GPT-5.5
78.65 ± 1.59
88.76
–
–
–
Opus-4.8
77.53 ± 2.43
88.76
–
–
–
DeepSeek-V4-Flash
77.15 ± 1.91
87.64
–
81.33 ± 0.50
–
GLM-5.2 (Teacher)
78.65 ± 1.84
86.52
–
83.33 ± 2.49
–
Qwen3.5-9B
26.59 ± 2.31
37.08
–
62.00 ± 0.28
–
Table 2: Main results on TerminalBench2.1 and the out-of-distribution SWE-bench Verified benchmark. Pass@1 is the mean ± standard deviation over three runs, and Δ denotes the absolute Pass@1 difference from Base.
Figure 3: Decomposition of final accuracy. Blue denotes initially correct candidates that remain correct, green denotes successful transitions from an incorrect candidate to a correct final solution, and red denotes harmful transitions from a correct candidate to an incorrect final solution.
Prefix generator
Pass@1 ↑
Qwen3.5-9B
33.33 ± 1.40
Qwen3.5-27B
31.46 ± 4.21
Qwen3.5-35B-A3B
31.83 ± 5.05
GLM-5.2 (teacher)
29.21 ± 2.42
Table 3: Effect of prefix source on Qwen3.5-9B.
Figure 4: Performance versus inference cost on TerminalBench2.1. Costs are averaged over common-success tasks within each backbone.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Candidate formed
Replay succeeded
Qwen3.5-9B
241/267 (90.26%)
233/240 (97.08%)
Qwen3.5-27B
263/267 (98.50%)
254/261 (97.32%)
Qwen3.5-35B-A3B
261/267 (97.75%)
250/259 (96.53%)
Qwen3.5-122B-A10B
261/267 (97.75%)
254/260 (97.69%)
Qwen3.6-27B
226/267 (84.64%)
220/225 (97.78%)
Qwen3.6-35B-A3B
256/267 (95.88%)
245/250 (98.00%)
Appendix
Table 4: Coverage of candidate formation and candidate-state replay on TerminalBench2.1. Entries report count/denominator (percentage), pooled across three independent runs. One Opus-4.8 trajectory was unavailable.
Model
Initial Acc.
Final Acc.
VTR
VOR
ESP
EDR
CAR
VPC
RSR
Qwen3.5-9B
12.02
26.60
99.59
99.58
98.69
69.75
92.59
29.75
23.65
Qwen3.5-27B
21.27
46.44
99.25
99.21
95.64
66.00
88.87
42.15
32.70
Qwen3.5-35B-A3B
19.93
38.20
99.23
99.18
95.97
71.53
87.41
43.85
25.35
Qwen3.5-122B-A10B
24.81
47.94
99.62
100.00
95.12
62.84
90.48
44.71
36.11
Qwen3.6-27B
32.01
56.93
99.59
99.54
92.91
63.44
90.43
54.07
54.95
Qwen3.6-35B-A3B
25.33
45.32
98.05
98.78
94.84
70.57
86.75
50.56
30.62
Appendix
Table 5: Per-model self-verification diagnostics on TerminalBench2.1. All values are percentages averaged over three independent runs. Initial Acc. (ICA) and Final Acc. denote accuracy at candidate formation and at the end of the trajectory, respectively; the remaining metrics are defined in Section 2.2 .
TerminalBench2.1
SWE-bench Verified (OOD)
Model
Run 1
Run 2
Run 3
Run 1
Run 2
Run 3
GPT-5.5
80.90
77.53
77.53
–
–
–
Opus-4.8
80.90
76.40
75.28
–
–
–
DeepSeek-V4-Flash
75.28
76.40
79.78
81.20
80.80
82.00
GLM-5.2 (Teacher)
80.90
78.65
76.40
86.00
84.00
80.00
Qwen3.5-9B
23.60
26.97
29.21
61.60
62.20
62.20
Appendix
Table 6: Per-run Pass@1 results for the main experiments. Each column reports an independent evaluation run, and all values are percentages. Table 2 reports the corresponding aggregate statistics.
Backbone
Method
N
Pass@1
Turns
Tool calls
Input tok.
Generated tok.
Total tok.
9B
Base
27
26.59±2.31
31.47
58.23
667.75
14.90
682.65
9B
FTD
27
27.72±3.71
19.67
48.75
539.36
54.20
593.56
9B
SCVD
27
36.33±2.12
20.28
48.27
468.79
61.01
529.80
27B
Base
43
46.44±0.53
23.91
42.41
528.22
16.24
544.47
27B
FTD
43
55.06±4.00
14.85
32.71
309.10
35.86
344.96
27B
SCVD
43
63.30±2.12
17.93
38.67
396.61
42.15
438.77
Appendix
Table 7: Detailed performance and inference-cost statistics on TerminalBench2.1. Pass@1 is evaluated over all 89 tasks, whereas inference costs are measured on common-success tasks within each backbone.
Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verification would unlock both, but the capability we want to train, i.e., catching self-generated errors, lacks training signal. To address this challenge, we propose self-trained verification (STV). Our key observation is that, while a model cannot catch these errors alone, it can when shown the reference solution. We turn this asymmetry into a supervision target and train the verifier to imitate a more informed version of itself. At test time, STV substantially improves V-R loops on hard problems, while alternatives (e.g., SFT, RL on verifier scores, and even meta-verifiers) do not. STV roughly doubles accuracy on hard math and lifts it 14x on scientific reasoning tasks (1.5% to 21%). At training time, we additionally train the generator using RL with STV verifier's feedback inside the V-R loop - a procedure we call verifier-in-the-loop training (ViL). Starting from an RL-converged generator, ViL yields a further 33% gain in pass@1. More notably, the generator's standalone pass@1, with no verifier at test time, climbs 30% relative past where standard RL had converged. Hence, the next frontier in reasoning on hard problems may lie in how we train for and with verification. Website: https://ar-forum.github.io/stv-webpage
Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit. We present SEVA, a structured verification agent that emits evidence alignments, step-by-step reasoning chains, calibrated confidence, and a six-category error diagnosis with actionable fixes. Training such an agent with RL is non-trivial: standard binary reward on multi-component output triggers advantage collapse -- within-group reward variance vanishes and the GRPO gradient disappears. We resolve this with a process reward that decomposes verification quality into five independent components weighted 70/30 toward process signals, restoring the gradient and inducing an implicit curriculum -- the agent first masters verification behavior (alignment 0.917 -> 0.997, format 72% -> 100%), then outcomes (F1 64.9 -> 69.0). Structured output further enables a Verify -> Reflect -> Probe -> Refine self-evolution loop, which over four rounds on a 7B model surfaces an unexpected structural finding: each round produces a benchmark-specialist, not a generalist (+15 pp on HaluEval, -10 to -14 pp on TruthfulQA in the same model, persistent at 4x data). On ClearFacts, SEVA-3B matches GPT-4o-mini (69.0 vs. 69.8 F1) while producing substantially richer, auditable output -- confirming a principle that should generalize: for any RL task with multi-component generation, reward granularity must match output granularity.
Aojie Yuan, Yi Nian, Haiyue Zhang +2
University of Southern California, Los Angeles, USA · University of Michigan, Ann Arbor, USA
Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.
Diandian Guo, Cong Cao, Fangfang Yuan +3
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China