Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.
Figures & tables
Figure 1: Endpoint correctness does not guarantee a valid evidence-to-action trace. An action may be permitted and succeed with the correct arguments, yet remain unsupported if required evidence was not established before execution. Missing support can propagate to downstream actions.
Figure 2: Overview of SafeActBench . A static decision setting and four interactive regimes span investigated non-action to dependency-constrained execution. A provenance-bound Evidence Ledger and deterministic evaluator verify required evidence, exact actions, and dependencies.
Protocol
Behavior
Actions
Strict success condition
# Cases
Static assessment (decision only, no trajectory scoring)
Legacy
Static decision
–
Correct Allow / Block / Defer judgment on a fixed candidate action
86
Interactive execution (full trajectory scored by the deterministic evaluator)
V0
Investigated non-action
0
Required investigation is completed and no consequential action occurs
175
V1
Single action
1
Required evidence is established before exactly one correct action
131
V2
Linear multi-action
≥ 2, chain
Evidence precedes each action, and later actions use actual predecessor results
132
Table 1: SafeActBench protocols , from static judgment to dependency-constrained execution.
Overall
Protocols
Model
Harness
ECS
P-M
D-M
Legacy
V0
V1
V2
V3
Claude Code
67.2
69.6
68.1
97.7
63.4
60.3
60.6
65.9
Claude-5
Inspect
63.4
66.0
65.4
96.5
58.9
54.2
59.1
61.4
Codex
65.4
67.5
67.8
96.5
65.7
52.7
58.3
64.4
GPT-5.6
Inspect
61.0
63.4
62.4
94.2
59.4
47.3
55.3
60.6
DSH
63.0
66.2
63.4
96.5
49.1
61.1
65.2
59.1
Table 2: Main results on SafeActBench across ten model–harness configurations. All values are percentages. ECS: exact case success. P-M and D-M: unweighted protocol- and domain-macro averages. Bold and underline indicate the best and second-best value in each column, respectively.
V0
V1
V2
V3
Model
Harness
BSR ↓
PAR ↓
CAS ↑
Gap ↓
Part. ↓
Gap ↓
Part. ↓
Claude-5
Claude Code
34.9
38.8
100.0
36.4
27.3
24.2
16.7
Inspect
37.1
43.2
100.0
37.9
27.3
26.5
18.2
GPT-5.6
Codex
29.7
45.8
97.2
29.5
20.5
27.3
21.2
Inspect
35.4
50.4
95.4
31.1
22.0
30.3
23.5
DeepSeek-V4
DSH
35.4
38.9
100.0
25.0
21.2
24.2
24.2
Table 3: Failure diagnostics on SafeActBench (%). BSR: stopped before required investigation (V0). PAR: acted before evidence completion (V1). CAS: success given an action after evidence completion (V1). Gap: unresolved requirement at an action checkpoint. Part.: executed but incomplete workflow (V2/V3). Darker shading marks worse values (definitions in Appendix B.6 ).
Paired Success
Inv. Comp.
Model
Fam.
Inspect
Δ [95% CI]
Fam./Insp.
Claude-5
62.6
58.4
+4.2 [ −0.7,+9.1 ]
64.4 / 60.5
GPT-5.6
60.7
56.0
+4.7 [ 0.0,+9.5 ]
62.6 / 57.7
DeepSeek-V4
57.9
53.5
+4.4 [ +1.1,+7.9 ]
58.9 / 59.5
Qwen3.8
62.1
59.3
+2.8 [ −1.2,+6.8 ]
63.2 / 60.9
GLM-5.2
28.6
35.4
−6.8 [ −11.2,−2.5 ]
29.8 / 38.8
Table 4: Paired harness comparison on the same 570 V0–V3 cases per model. Fam.: family-associated harness. Δ : Fam. minus Inspect in percentage points, with 95% case-paired bootstrap CIs (bold when the CI excludes zero). Inv. Comp.: investigation-completion rate. Rates are percentages.
Figure 3: Evidence sensitivity under controlled interventions. (a) Action probability among completed episodes per condition. (b) Paired reductions in action probability relative to Original on jointly completed cases. Points and lines give estimates and 95% CIs from 10,000 paired case-bootstrap resamples. Pooling details, sample sizes, and coverage are given in Appendix C.4 .
Figure 4: Mechanism probes under missing evidence. (a) Action probability under the Withheld, evidence-package, and requester-claim conditions. (b) Paired changes in action probability with 95% paired case-bootstrap confidence intervals. Distractors are compared with the evidence package and all other probes with Withheld. (c) Mean information calls with the reference evidence package and with added distractors, under Original and Withheld evidence. Matched sample sizes are 43 cases per configuration, except Qwen urgency in (b) and GLM distractors in (b,c), which use 42.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Work
Stateful
Trajectory Checks
Pre-Action Evidence
Entity/State Binding
Result Propagation
Investigated Non-Action
τ -bench ( Yao et al., 2025 )
✓
–
–
–
–
–
ToolSandbox ( Lu et al., 2025 )
✓
✓
△
△
✓
△
AppWorld ( Trivedi et al., 2024 )
✓
–
–
–
–
–
Agent-SafetyBench ( Zhang et al., 2024 )
✓
✓
△
–
–
△
Near-Miss ( Rabinovich et al., 2026 )
✓
✓
✓
△
–
–
EnvTrustBench ( Sheng et al., 2026 )
✓
✓
△
△
–
△
Appendix
Table 5: Comparison of selected benchmarks and evaluation methods for tool-using agents. Stateful denotes support for persistent environment state. For evaluation properties, ✓ denotes an explicit scoring requirement, △ denotes partial or task-specific coverage, and – denotes no explicit criterion. Entries describe evaluation protocols rather than behaviors that may arise incidentally while solving a task.
Domain
Legacy
V0
V1
V2
V3
Total
Customer and policy operations
14
29
23
22
24
112
Engineering and infrastructure operations
13
28
23
22
23
109
Legal and financial operations
15
42
31
27
29
144
Research assistance
14
24
20
21
18
97
Smart-home control
15
25
17
20
19
96
Healthcare operations
15
27
17
20
19
98
Appendix
Table 6: Distribution of SafeActBench cases across protocols and operational domains.
Investigation
Workflow Structure
Protocol
Mean
Median
Max.
Property
V2
V3
V0
5.59
6
15
Action nodes
317
501
V1
6.47
6
10
Dependency edges
185
439
V2
7.14
7
12
Result references
182
296
V3
5.88
5
17
Cases with a fork
0
103
Cases with a join
0
96
Appendix
Table 7: Investigation and workflow complexity in SafeActBench . The left panel reports required tool–argument reads per case, excluding additional discovery queries. The right panel summarizes the aggregate structure of V2 and V3. Result references count occurrences in downstream action arguments, while fork and join counts denote cases containing each structure.
Variation
Accepted / tested
Additional information read
395/395
Independent V3 action swap
132/132
Equivalent numeric representation
24/24
Alternative retrieval path
32/32
Equivalent facts from different tools
89/89
Joint support from multiple observations
25/25
Appendix
Table 8: Evaluator behavior on valid trajectory variations, alternative evidence paths, and negative controls. Counts report accepted/tested variants after evaluator repair.
Model
Harness
Static accuracy
Interactive ECS
Gap [95% CI]
DeepSeek
Inspect
95.0
52.0
+43.0[32.0,54.0]
GLM
Inspect
97.0
28.0
+69.0[59.0,78.0]
GLM
ZCode
99.0
33.0
+66.0[56.0,75.0]
Appendix
Table 9: Static judgment and interactive execution on the same 100 V1 cases per configuration. Static accuracy measures correct judgments of valid supplied actions with complete evidence, whereas Interactive ECS measures success on the original V1 tasks. Rates are percentages. The paired gap is Static minus Interactive, expressed in percentage points with a 95% confidence interval. The 24 additional negative examples per configuration are evaluated separately and are excluded from these denominators.
Model
Harness
Scored S/I
Paired N
Static
Interactive
Δ [95% CI]
Claude-5
Claude Code
99/99
99
100.0
94.9
+5.1 [ +1.0,+9.1 ]
Claude-5
Inspect
99/99
99
94.9
72.7
+22.2 [ +12.1,+32.3 ]
GPT-5.6
Codex
99/99
99
99.0
93.9
+5.1 [ +1.0,+9.1 ]
GPT-5.6
Inspect
99/99
99
97.0
90.9
+6.1 [ +2.0,+11.1 ]
DeepSeek-V4
DSH
99/95
87
93.1
85.1
+8.0 [ +1.1,+16.1 ]
DeepSeek-V4
Inspect
98/94
81
90.1
79.0
+11.1 [ +1.2,+21.0 ]
Appendix
Table 10: Balanced static versus interactive decision accuracy. Scored S/I gives scored static and interactive episodes out of 99 planned per condition. Paired N counts scenarios from complete base-task triples, so the number of independent base tasks is N/3 . Accuracy is computed on these identical, class-balanced paired subsets and is expressed as a percentage. Δ is static minus interactive accuracy in percentage points, with 95% base-task cluster-bootstrap confidence intervals. Invalid outputs remain incorrect, and unscored episodes are excluded from paired estimates but remain visible in coverage. Dashes indicate insufficient independent paired coverage, not zero accuracy.
Figure 5: Failure modes across SafeActBench protocols. All values are proportions of all cases within each model–harness configuration and protocol, rounded to two decimal places. In V1, early action denotes acting before evidence completion, while covered action failure denotes a failed action after evidence completion. In V2 and V3, gap denotes an unresolved requirement at an action checkpoint, and partial denotes executed actions with incomplete workflow execution. Gap only, partial only, and gap + partial are mutually exclusive. Other failure includes the remaining scored failures. Total failure sums all scored-failure categories before rounding and excludes unscored cases.
Completed / planned (coverage)
Matched pairs
Configuration
Original
Withheld
Contradicted
Original–W
Original–C
DeepSeek–DSH
85/86 (98.8%)
43/43 (100.0%)
21/22 (95.5%)
42
20
GLM–ZCode
86/86 (100.0%)
43/43 (100.0%)
22/22 (100.0%)
43
22
Qwen–Qwen Code
86/86 (100.0%)
43/43 (100.0%)
22/22 (100.0%)
43
22
Appendix
Table 11: Completion coverage and matched case counts for the base evidence interventions. Completion cells give completed/planned episodes and coverage. Original completion pools two replicates, whereas matched pairs use Original replicate 1. W and C denote Withheld and Contradicted.
Figure 6: Additional evidence-intervention results. (a) Descriptive action rates over completed episodes. (b) Paired reductions relative to Original replicate 1, with 95% paired case-bootstrap confidence intervals. (c) Matched decision transitions, retaining invalid outputs. (d) Mean queries to the affected information tool, and fractions report Withheld episodes with an action that queried that tool. Completion counts and pairing denominators appear in Table 11 .
Figure 7: Investigation effort and decision readiness in historical trajectories from five models using their family-associated harnesses. Columns correspond to V0–V3. (A) Fraction of scored episodes with complete investigation at the recorded endpoint and normalized query cost at most the horizontal-axis value. The horizontal axis is linear up to one and logarithmic thereafter. (B) Cumulative distributions of required-read coverage at stopping (V0), coverage before the first action (V1), and case-weighted local requirement support at action gates (V2/V3). Greater cumulative mass below one indicates more decisions with incomplete coverage or support. Shading shows pointwise 95% confidence intervals from 10,000 case-bootstrap resamples, retaining all gates of a sampled case together. All five models cover the same 570 cases, but these historical runs are separate from the main-results evaluation.
Figure 8: First-action timing under additional interventions. Columns correspond to DeepSeek–DSH, GLM–ZCode, and Qwen–Qwen Code. Panels (a)–(c) compare Withheld with evidence-package and requester-claim conditions, and panels (d)–(f) compare Withheld with urgency and explicit-warning conditions. Each panel uses cases completed under all three displayed conditions: 43 cases per panel, except panel (f), which uses 42. Consequently, Qwen’s Withheld baseline is 23/43 in panel (c) and 23/42 in panel (f). Colored annotations report final action rates. Shading denotes pointwise 95% paired case-bootstrap confidence intervals, using the resampling procedure described in Appendix C.4 .
Figure 9: Retrieval effort with and without distractor records. Curves show the cumulative share of matched completed cases using at most k information calls, and a rightward shift indicates greater retrieval effort. Columns correspond to DeepSeek–DSH, GLM–ZCode, and Qwen–Qwen Code. Panels (a)–(c) use Original evidence, and panels (d)–(f) use Withheld evidence. Each panel compares the reference evidence package with its distractor-augmented counterpart. Matched sample sizes are 43 for DeepSeek, 42 for GLM, and 43 for Qwen in both rows. Annotations report paired changes in mean information calls and action probability, computed as distractor minus reference. Action-probability changes are expressed in percentage points. Shading shows pointwise 95% paired case-bootstrap confidence intervals for the cumulative curves, and bracketed annotations give 95% paired confidence intervals for the differences, using the procedure described in Appendix C.4 .
Figure 10: Evidence sensitivity across native and Inspect harnesses. Native denotes DSH for DeepSeek and ZCode for GLM. W and C denote Withheld and Contradicted. (a) Paired action-probability reductions relative to Original replicate 1. (b) Inspect-minus-native differences in those reductions. Matched sample sizes for W/C are 42/20 for DeepSeek and 43/22 for GLM. Different intervention subsets preclude directly ranking W and C from panel (a). (c,d) Action probabilities on common six-cell probe subsets: 36 DeepSeek cases and 43 GLM cases. All values are proportions, and error bars show pointwise 95% case-bootstrap intervals from 10,000 joint resamples. Missing episodes are excluded, and invalid outputs are not counted as valid refusals.
Model
Quantity
R1
R2
R3
Mean ± SD
Claude-5
Fam.
63.0
62.1
62.8
62.6±0.5
Inspect
58.2
58.4
58.6
58.4±0.2
Δ
4.7
3.7
4.2
4.2±0.5
GPT-5.6
Fam.
61.1
60.4
60.7
60.7±0.4
Inspect
55.6
56.1
56.1
56.0±0.3
Δ
5.4
4.2
4.6
4.7±0.6
Appendix
Table 12: Success rates and harness differences across three repetitions. Fam. denotes the family-associated harness. Fam. and Inspect rows report ECS in percent, and Δ rows report differences in percentage points. R1–R3 denote the three repetitions, and Mean ± SD reports their mean and sample standard deviation. All results use the same 570 V0–V3 cases, with 100% evaluation coverage in each repetition.
Figure 11: Overview of SCGR. SCGR organizes model-selected investigation, evidence reuse, and public checks before submission. The model selects questions and queries, assesses the evidence, and revises its candidate. The runtime records observations and checks public structural constraints but does not determine evidence sufficiency.
Overall
Protocols
Model
Harness
Setting
ECS
P-M
D-M
Legacy
V0
V1
V2
V3
Baseline
67.2
69.6
68.1
97.7
63.4
60.3
60.6
65.9
Claude Code
+SCGR-EG
67.6 (+0.4)
69.9
68.4
97.7
64.2
60.8
60.4
66.3
Baseline
63.4
66.0
65.4
96.5
58.9
54.2
59.1
61.4
Claude-5
Inspect
+SCGR-EG
63.7 (+0.3)
66.3
65.7
96.5
59.5
54.8
58.7
61.8
Baseline
65.4
67.5
67.8
96.5
65.7
52.7
58.3
64.4
Appendix
Table 13: Baseline and SCGR-EG results on SafeActBench across ten model–harness configurations. Baseline rows and metrics follow Table 2 , and all values are percentages. Parenthesized values give the ECS change over the baseline in points.
Model
Harness
Method
N
Run Comp.
ECS
V0
V1
V2
V3
Full V0–V3 coverage: n=(175,131,132,132)
Claude-5
Claude Code
Baseline
570
100.0
62.6
63.4
60.3
60.6
65.9
SCGR-Select
570
100.0
53.5
58.9
66.4
47.7
39.4
GPT-5.6
Codex
Baseline
570
100.0
60.7
65.7
52.7
58.3
64.4
SCGR-Select
570
88.8
58.1
65.7
74.0
57.6
32.6
Qwen3.8
Qwen Code
Baseline
570
100.0
62.1
72.6
58.8
59.8
53.8
Appendix
Table 14: Structured evidence selection as an analysis intervention. Baseline rows follow the main evaluation (Table 2 ), with ECS computed over the included V0–V3 cases. N denotes the number of cases, and Run Comp. denotes run completion rate. All rates are percentages. Full coverage includes all 570 V0–V3 cases. DeepSeek SCGR-Select results use partial cohorts after excluding SCGR-EG entries, so the two DeepSeek–DSH rows cover different case sets. Legacy is excluded, and dashes indicate protocols not included in the comparison.
An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero for clock- or deadline-driven stopping. Where we record their judgments, the seven agents we test call a failing source's results useless 97-100% of the time, yet most of them rarely stop on that judgment. Prompt cues change when they stop but not what they stop on. Permission to answer from memory and a reasoning mode can bring early stops regardless of evidence, a stated budget moves the 7-8B models' stops to the deadline, and a stopping rule or call cost in the prompt is followed at most partly. Stopping follows the evidence only when the harness enforces an integration step that makes the agent answer after five consecutive results it judged useless. This step raises failing-source success for every model, keeps the stopping point fixed when the budget doubles, and needs no extra judgment call when the agent states its judgments. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.
Chubin Zhang, Zhenglin Wan, Xingrui Yu +4
Nanyang Technological University, Singapore · National University of Singapore, Singapore · CFAR, Agency for Science, Technology and Research, Singapore +3
Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately adopted. We introduce ToxicBench to measure checking and adoption under numerical, label, schema, and retrieval errors, pairing clean and poisoned observations over fixed source data. In the 118-task GPT evaluation across three adapters, poisoning lowers task success by 26 to 39 percentage points. Ordinary retries help under one-shot poisoning, whereas repeated poisoning reveals wrong-answer adoption after checking. Controls on three public tables isolate how supplied evidence affects recovery. After freezing the scorer, we compare its judgments with human annotations on 200 trajectories, finding 96% task-success agreement. Human judgments support retry gains over Base and confirm adoption after checking on audited tasks. We release trajectories, versioned scoring, and reference and delivery audits. These findings highlight evidence availability and answer selection as complementary dimensions of agent reliability.
Tool-using agents increasingly rely on external tools to complete multi-step tasks, but tool returns can fail in different ways and require different recovery actions. Existing robustness studies often use uncertainty-based measures to detect when an agent becomes unreliable. These measures can reveal that something has gone wrong, but they do not directly identify the type of tool failure or the appropriate response. We address this limitation by analyzing tool failures at the moment a return enters the agent context. Our approach combines two complementary signals. The first compares the likelihood of the returned content under the tool schema and under the full trajectory prefix. The second measures the agent's probability distribution over its legal next actions. We evaluate the approach by injecting incomplete and inconsistent returns into a retail customer-service benchmark. The results show that likelihood-based signals clearly capture incomplete returns and some direct inconsistencies, while action-based signals reveal how strongly a failure changes the next decision. Some failures that are weak under likelihood signals can still redirect the agent toward state-changing actions. These findings show that tool failures can be recognized at the return boundary, but reliable diagnosis requires combining multiple signals.
Jiachen Xu, Torben Bach Pedersen, Zhongming Yao +2
Department of Computer Science Aalborg University Aalborg, Denmark · College of Computer Science and Technology Zhejiang University Hangzhou, China · School of Information Science and Engineering Northeastern University Shenyang, China