Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification
Organizations: Northeastern University · Inclusion AI, Ant Group · University of Maryland, College Park
Abstract
Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifies the first complete solution in each trajectory, determines whether it is objectively correct, and uses this ground truth to quantify the agent's subsequent verification and recovery behavior. Applying it to ten terminal agents on TerminalBench2.1, we find that verification is nearly universal after a complete candidate is formed, yet only 61.43% of incorrect candidates are detected and only 49.36% of detected errors are successfully repaired. These results show that the main weakness in self-verification lies not in initiating verification, but in detecting and repairing errors. Motivated by these findings, we propose Student-Conditioned Verification Distillation (SCVD), which lets the student first produce a candidate solution and distills a stronger teacher's subsequent verification and recovery from the same interaction context. Across three Qwen3.5 backbones, SCVD improves \textsc{Pass@1} on TerminalBench2.1 by 9.74--16.85 percentage points over the corresponding base models and by 4.49--8.61 points over the standard full-trajectory distillation, while avoiding the pronounced out-of-distribution degradation of full-trajectory distillation on SWE-bench Verified.
Figures & tables
| Candidate correctness | |||
|---|---|---|---|
| Incorrect ( ) | Detected error | Missed error | No usable outcome |
| Correct ( ) | False alarm | Correct pass | No usable outcome |
| TerminalBench2.1 | SWE-bench Verified (OOD) | |||||
| Model | Pass@1 | Pass@3 | Pass@1 | |||
| GPT-5.5 | 78.65 1.59 | 88.76 | – | – | – | |
| Opus-4.8 | 77.53 2.43 | 88.76 | – | – | – | |
| DeepSeek-V4-Flash | 77.15 1.91 | 87.64 | – | 81.33 0.50 | – | |
| GLM-5.2 (Teacher) | 78.65 1.84 | 86.52 | – | 83.33 2.49 | – | |
| Qwen3.5-9B | 26.59 2.31 | 37.08 | – | 62.00 0.28 | – | |
| Prefix generator | Pass@1 |
|---|---|
| Qwen3.5-9B | 33.33 1.40 |
| Qwen3.5-27B | 31.46 4.21 |
| Qwen3.5-35B-A3B | 31.83 5.05 |
| GLM-5.2 (teacher) | 29.21 2.42 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Candidate formed | Replay succeeded |
|---|---|---|
| Qwen3.5-9B | 241/267 (90.26%) | 233/240 (97.08%) |
| Qwen3.5-27B | 263/267 (98.50%) | 254/261 (97.32%) |
| Qwen3.5-35B-A3B | 261/267 (97.75%) | 250/259 (96.53%) |
| Qwen3.5-122B-A10B | 261/267 (97.75%) | 254/260 (97.69%) |
| Qwen3.6-27B | 226/267 (84.64%) | 220/225 (97.78%) |
| Qwen3.6-35B-A3B | 256/267 (95.88%) | 245/250 (98.00%) |
| Model | Initial Acc. | Final Acc. | VTR | VOR | ESP | EDR | CAR | VPC | RSR |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | 12.02 | 26.60 | 99.59 | 99.58 | 98.69 | 69.75 | 92.59 | 29.75 | 23.65 |
| Qwen3.5-27B | 21.27 | 46.44 | 99.25 | 99.21 | 95.64 | 66.00 | 88.87 | 42.15 | 32.70 |
| Qwen3.5-35B-A3B | 19.93 | 38.20 | 99.23 | 99.18 | 95.97 | 71.53 | 87.41 | 43.85 | 25.35 |
| Qwen3.5-122B-A10B | 24.81 | 47.94 | 99.62 | 100.00 | 95.12 | 62.84 | 90.48 | 44.71 | 36.11 |
| Qwen3.6-27B | 32.01 | 56.93 | 99.59 | 99.54 | 92.91 | 63.44 | 90.43 | 54.07 | 54.95 |
| Qwen3.6-35B-A3B | 25.33 | 45.32 | 98.05 | 98.78 | 94.84 | 70.57 | 86.75 | 50.56 | 30.62 |
| TerminalBench2.1 | SWE-bench Verified (OOD) | ||||||
| Model | Run 1 | Run 2 | Run 3 | Run 1 | Run 2 | Run 3 | |
| GPT-5.5 | 80.90 | 77.53 | 77.53 | – | – | – | |
| Opus-4.8 | 80.90 | 76.40 | 75.28 | – | – | – | |
| DeepSeek-V4-Flash | 75.28 | 76.40 | 79.78 | 81.20 | 80.80 | 82.00 | |
| GLM-5.2 (Teacher) | 80.90 | 78.65 | 76.40 | 86.00 | 84.00 | 80.00 | |
| Qwen3.5-9B | 23.60 | 26.97 | 29.21 | 61.60 | 62.20 | 62.20 | |
| Backbone | Method | Pass@1 | Turns | Tool calls | Input tok. | Generated tok. | Total tok. | |
|---|---|---|---|---|---|---|---|---|
| 9B | Base | 27 | 31.47 | 58.23 | 667.75 | 14.90 | 682.65 | |
| 9B | FTD | 27 | 19.67 | 48.75 | 539.36 | 54.20 | 593.56 | |
| 9B | SCVD | 27 | 20.28 | 48.27 | 468.79 | 61.01 | 529.80 | |
| 27B | Base | 43 | 23.91 | 42.41 | 528.22 | 16.24 | 544.47 | |
| 27B | FTD | 43 | 14.85 | 32.71 | 309.10 | 35.86 | 344.96 | |
| 27B | SCVD | 43 | 17.93 | 38.67 | 396.61 | 42.15 | 438.77 |
| Hyperparameter | Value |
|---|---|
| Epochs | 3 |
| Optimizer | AdamW |
| Learning rate | |
| LR scheduler | Cosine |
| Warmup ratio | 0.03 |
| Maximum sequence length | 32,768 |