Predicting whether an LLM agent will fail has emerged as a promising direction for supporting intervention during execution. Recent approaches report strong predictive performance, often with AUROC values between 0.85 and 0.94. However, predictors are typically trained by pooling runs from many tasks. We hypothesize that part of this performance comes from recognizing that some tasks are harder than others, rather than detecting whether a particular run is heading toward failure. This distinction matters because task-level difficulty supports decisions about where to allocate computation, while run-level prediction is needed to decide whether to intervene in an ongoing trajectory. We study benchmarks with repeated attempts of the same task by the same agent and separate cross-unit comparisons from comparisons between successful and failed runs of the same model-task unit. Across the evaluated corpora, more than 99.93% of the positive-negative pairs underlying pooled AUROC are cross-unit. Accordingly, predictors that never observe the current run can achieve high pooled performance, including a difficulty oracle with AUROC up to 0.945, while remaining at chance within task. Early run-level discrimination is consistently weak across trajectory predictors, released monitors, and hidden-state probes, although it improves later in execution and is stronger for weaker agents. Under fixed token budgets, task-level allocation outperforms abort-only strategies, while early stopping becomes beneficial only when within-task AUROC reaches about 0.84-0.93, far above the 0.50-0.55 range observed for early monitors. These results show that failure prediction should be evaluated not only by outcome accuracy, but by whether the captured signal supports the intended deployment decision.
Figures & tables
Figure 1. Overview of the study design. Repeated attempts within fixed model-task units allow pooled AUROC to be separated into same-unit ( Awithin ) and cross-unit ( Across ) discrimination. We use this distinction to examine what pooled failure prediction measures, when run specific signal emerges, and how the resulting information supports deployment decisions.
C1
C2
C3
C4-L
C4-Q
Source
SWE-agent ( Yang et al., 2024 )
SWE-rebench ( Badertdinov et al., 2025 )
LiveClawBench ( Long et al., 2026 )
Latent Programming Horizons ( Silva et al., 2026 )
Scaffold
SWE-agent
OpenHands
OpenClaw
–
–
Trajectories
80,036
67,074
6,834
5,000
5,000
Tasks
3,591
6,306
134
500
500
Model(s)
3 (Llama-3 8B–405B)
Qwen3-Coder-480B
17
Laguna-XS.2
Qwen3.6-35B-A3B
Attempts/task
–
–
–
10
10
Table 1. Corpora used in the study. C1–C3 support the main behavioral analyses, while C4-L and C4-Q support the hidden-state analyses. Mixed-outcome is the share of units containing both a successful and a failed run.
Condition
n
Apool
Awithin
gap
units
k=2
79,929
0.6435
0.5169
+0.127
926
k=5
77,951
0.6863
0.5428
+0.144
920
k=10
63,657
0.7219
0.5987
+0.123
885
k=20
35,554
0.7463
0.6527
+0.094
620
Difficulty oracle
80,036
0.9454
0.5000
+0.445
926
Table 2. C1 results at fixed-turn checkpoints (Section 4.6 ). n is the number of evaluated trajectories, Apool and Awithin are pooled and within-task AUROC, gap is their difference, and units is the number of mixed model–task pairs.
Figure 2. Pooled versus within-task AUROC across representative prediction conditions. Run-blind predictors can achieve high pooled discrimination while remaining at chance within task, whereas predictors using current-run information show substantially weaker within-task discrimination. The shaded band marks the approximate within-task AUROC required for early aborting to improve throughput on C4-L.
Figure 3. Length baseline versus activation probe on C4-L.
Figure 4. Tasks solved on the 250-task C4-L split under fixed token budgets. Points average 150 paired replay trials; bars show 95% percentile intervals. Abort thresholds are tuned separately. TS denotes budgeted Thompson sampling with task information transferred from C4-Q.
LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself. We ask how much detection is achievable from observable step telemetry alone, using monitors costing microseconds per step and trained only on healthy runs. On 2,823 committed agent episodes across three frameworks, three local models (qwen2.5 7b/3b, llama3.1 8b) and a commercial API (gemini-2.5-flash), a one-class echo-state-network ensemble with CUSUM alarms detects 0.71 of failures at a 5% false-alarm budget (AUROC 0.872). Its advantage over a memoryless baseline is a monotone function of post-onset horizon (+0.09 at <=3 steps, +0.40 at >=9), predicting its own failure region out of sample on AFTraj-2K. Ranking transfers with no retraining to two corpora from other groups (AFTraj-2K 0.745, ATBench 0.779). Monitors carry two burdens: a per-deployment healthy null (they do not transfer -- AUROC 0.527 cold against 0.885 recalibrated) and a residual false-alarm rate. We add a layer carrying neither: deterministic verification, which recomputes a run's stated total from the tool results it actually received and confirms every required call was made. Head-to-head it catches 60% of failures (96% with the coverage check) at 0 of 63 false positives against the monitor's 54% at 17%, transfers unchanged to llama3.1:8b (110 of 110 at 0 of 10), and trips on 0 of 1825 healthy episodes. Detection is then closed into repair: each flagged run is rolled back and re-run live, recovering 45% of failures against a 16% resampling control (p=0.0005) and lifting task success from 52% to 73% for about one extra model call per run. The system runs at ~200 microseconds per step, three orders of magnitude below a judge call. Code, traces and results are released.
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure that records which action wins without reflecting the absolute level of performance. When all actions fail, they tie at zero reward, and independent runs produce the same tied set with high probability, creating an illusion of stability that masks near-zero recovery success. We formalize this limitation through a set-path symmetry result, proving that for equal-cost Bernoulli actions the success probabilities (0.9, 0.8) and (0.2, 0.1) yield identical best-action-set distributions at every sample size. No procedure based solely on which action wins can distinguish these two regimes. We further prove that certifying exact population ties is impossible in finite time, and that the assignment of outcomes to checkpoints carries information beyond marginal outcome distributions. The pooled success probability is the missing scalar that resolves the ordinal ambiguity. Experiments on 864 frozen RecoveryBench episodes and two planning cohorts totaling 3,456 responses confirm the theoretical predictions. Agreement and held-out quality can move in opposite directions, and permuting checkpoint-to-action bindings changes 8 to 13 percent of cell-level conclusions. Based on these findings, we propose reporting four diagnostic quantities (agreement, all-zero fraction, held-out success, and pooled success) that expose this failure mode with no additional data collection.
Dong Xu, Zhangfan Yang, Jiantao Wu +5
School of Artificial Intelligence, Shenzhen University · EasternDawn · School of Computer Science, University of Nottingham Ningbo
LLM agents can fail silently by asserting task completion when the environment state shows otherwise. We study this failure mode, false success, across two agent benchmarks: 9,876 tau2-bench trajectories from 8 model families and 1,879 AppWorld trajectories from 4 model families with text-independent ground truth. False success is common but varies by setting: 45--48% of failures in single-control tau2-bench domains, 3% in dual-control telecom, and 75.8% among AppWorld self-assessing coding-agent trajectories with explicit status claims. LLM judges fail reliably: no configuration across 5 judges, 5 prompt strategies, and full task specifications exceeds AUROC 0.65 on tau2-bench, and the same judges reach only 0.54 AUROC on AppWorld API-call traces. Judges rely on surface completion proxies -- confident closing language in tau2-bench and coarse action-sequence volume in AppWorld -- rather than verified state changes. Lightweight TF-IDF detectors achieve task-disjoint AUROC 0.83 on tau2-bench and 0.95 on AppWorld, recovering 4--8x more false successes than the best judge at the same flag rate with 3,300x lower latency. These results suggest that production monitoring should use lightweight, domain-calibrated detectors as triage signals rather than relying on LLM judges as the primary monitor for false success.