Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across business scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Our experiments reveal that even the strongest proprietary and open-weight models show steep reliability drops: Claude Opus 5 falls from 66.50% pass@1 to 47.53% pass^20, and Kimi-K3 from 57.37% pass@1 to 17.60% pass^20. Moreover, many failed trials terminate cleanly after valid state-changing actions, so response- or tool-call-level signals poorly proxy end-to-end completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox (https://github.com/microsoft/thinkingbox) and Thinkingbox-bench (https://github.com/microsoft/thinkingbox-data).
Figures & tables
Figure 1: (A) ThinkingBox runs multi-turn interactions among a simulated user, an LLM agent, and isolated MCP-compatible tools, then evaluates terminal state, side effects, and dialogue with executable judges. More detailed pipeline can be found in Figure 2 (B) Across 507 tasks and 20 attempts per task, pass@20 is much higher than pass@1, while all-20 success is much lower. This reveals a discovery–reliability gap : a model may be able to discover at least one successful trajectory across multiple attempts (pass@20), yet fail to reproduce that success reliably across repeated attempts (all-20 success). Full results can be found in Table 4
Figure 2: Overview of Thinkingbox . The sandbox orchestrates the interaction among a simulated user, LLM agents, isolated domain tools, side-effect extraction, and executable judges. The same trajectory-level verdict supports benchmark evaluation, failure analysis, and agent training.
Benchmark
Primary domain
Tools/APIs
User dialogue
Stateful backend
Side-effect checks
MCP servers
SWE-bench ( Jimenez et al., 2024 )
Code repair
×
×
✓
×
×
BFCL ( Patil et al., 2025 )
Function calling
✓
×
×
×
×
ToolBench / API-Bank ( Qin et al., 2024 ; Li et al., 2023 )
API tool use
✓
×
×
×
×
WebArena / OSWorld ( Zhou et al., 2024 ; Xie et al., 2024 )
Web/desktop control
×
×
✓
×
×
AppWorld ( Trivedi et al., 2024 )
App APIs / coding agents
✓
×
✓
✓
×
MCP-Atlas ( Bandi et al., 2026 )
Real MCP servers
✓
×
×
×
✓
Table 1: Comparison with representative agent benchmarks. △ denotes partial or indirect support. Thinkingbox-bench focuses on stateful business-domain workflows where the final verdict checks backend state, side effects, and user dialogue, while exposing tasks through MCP servers.
Retail
Booking
Insurance
Neobank
Consulting
Tasks
98
104
100
104
101
Backend systems
11
8
7
3
18
Databases (tables / rows)
22 / 86
17 / 98
14 / 72
20 / 151
30 / 231
Agent tools (write / read)
16 / 17
10 / 28
14 / 19
13 / 19
13 / 14
Policy (words)
945
3,684
2,471
3,392
1,747
Knowledge base (documents)
9
11
8
8
9
Table 2: Key statistics for the ThinkingBox-Bench domains.
Domain
Scenario
Required agent behavior
Executable checks
Retail / e-commerce
User asks to change or refund part of an order.
Identify the correct order/item, verify eligibility, request missing confirmation, and update only the relevant record.
Correct order state; refund/order side effect; no unrelated customer or item modified.
Travel / hospitality
User requests a booking change under date, room, or policy constraints.
Check reservation, availability, and change policy before modifying booking or explaining denial.
Correct reservation state; price/fee side effect if applicable; policy-compliant dialogue.
Auto insurance
User reports or updates a claim.
Verify policy and vehicle/incident details, collect missing information, and create or update the claim.
Correct ticket/claim state; no coverage mutation unless allowed.
Neobank internal IT
Employee requests access to an internal application.
Verify employee role, existing access, and required approvals before provisioning or escalating.
Correct access, approval, and ticket state; no unauthorized privileges or unrelated records changed.
Consulting IT/HR
Employee requests enrollment in remaining onboarding courses.
Check existing enrollment, add missing courses, and keep the onboarding ticket pending until completion.
Correct course enrollment and ticket state; no duplicate enrollment or premature closure.
Table 3: Representative Thinkingbox-bench scenario patterns. Each scenario is instantiated as a concrete task with an initial backend state, available MCP-compatible tools, and executable checks.
Model
Size
Retail (98)
Auto (100)
Booking (104)
Bank (104)
Consulting (101)
Average
Proprietary models
Claude Opus 5
–
80.71 ± 3.68
65.80 ± 4.24
49.95 ± 4.38
70.62 ± 4.10
66.19 ± 4.32
66.50 ± 1.91
GPT-5.4
–
76.33 ± 3.53
62.65 ± 3.71
68.12 ± 3.46
65.34 ± 3.17
54.60 ± 4.00
65.36 ± 1.63
GPT-5.6-sol
–
67.65 ± 3.55
65.30 ± 3.66
60.34 ± 3.65
59.09 ± 3.64
57.52 ± 4.02
61.91 ± 1.66
Claude Sonnet 4.6
–
72.35 ± 3.65
54.40 ± 4.03
58.94 ± 3.8
56.39 ± 3.40
54.31 ± 3.73
59.19 ± 1.69
GPT-6 Astra
–
71.73 ± 4.38
46.55 ± 4.31
55.87 ± 4.68
60.87 ± 4.41
56.83 ± 4.56
58.31 ± 2.03
Table 4: Thinkingbox-bench pass@1 (%) by domain, micro-averaged over repeated N=20 trials per task. ± indicates standard error across tasks. Size denotes total/activated parameters for MoE models; bold/underline indicate best/second-best within each group. See Appendix C.4 for additional training details and Appendix D.1 for complete inference intervals.
Model
Tool Usage
No State-Changing Action
Incomplete User Resolution
Wrong State Update
Claude Opus 5
96.4
0.8
0.8
2.0
GPT-5.4
89.6
1.6
0.5
8.3
GPT-5.6-sol
86.6
8.6
0.2
4.6
Claude Sonnet 4.6
84.0
2.8
4.2
8.9
GPT-6 Astra
97.0
0.8
0.3
1.9
Claude Opus 4.6
78.1
3.0
10.2
8.7
Table 5: Failure mode breakdown (%) on Thinkingbox-bench . Values show each model’s distribution over dominant failure types; Appendix D.4 provides complete failure example trajectories.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Backend only
Backend + rubric
Task families represented in the final set
Retail / e-commerce
98
0
Delivery delay and exception handling; delivered-but-missing and return-to-sender cases; returns, refunds, exchanges, and warranty claims; installation scheduling and cancellation; order cancellation, promotions, and membership changes.
Travel / hospitality
89
15
Individual, corporate, and group booking changes; payment recovery; cancellations and refunds; corporate invoices and account benefits; group billing and services; hotel-partner verification and discrepancies; special requests and post-stay complaints.
Auto insurance
100
0
Billing extensions and arrangements; proof-of-insurance documents; adding, removing, or updating drivers and vehicles; first notice of loss and claim intake; listed-driver requests; reinstatement and policy cancellation.
Neobank internal IT support
89
15
Employee access and approval requests; password and account-security actions; production-incident access; hardware troubleshooting, assignment, replacement, and procurement; software and license requests; policy-information questions.
Consulting IT / HR support
101
0
Client-system and document access; software provisioning; expenses; hardware requests; employee onboarding; training enrollment; engagement and approval checks; corporate-travel policy and escalation.
Total
477
30
507 executable cases in total.
Appendix
Table 6: Composition of the 507-task evaluation set. Every case has executable backend-state checks. “Backend only” cases use no additional response rubric, whereas “Backend + rubric” cases also impose a binary final-response requirement.
Artifact
Contents
Review question
User goal g
Initial request plus facts the simulated user can provide during follow-up
Is the request natural, internally consistent, and resolvable without access to hidden evaluator information?
Initial state b0
Synthetic records in the domain backend, including existing tickets and related business objects
Do all referenced identifiers resolve, and do cross-system records agree before the agent acts?
Policy context
Domain operating manual and the fixed evaluation time
Does the policy determine eligibility, approvals, disclosures, and allowed actions without exposing the golden result?
Tools T
MCP-compatible read and write operations over the isolated domain services
Can the required evidence be retrieved and the intended outcome be executed using available tools?
User simulation U
Task-specific user role used for on-policy follow-up dialogue
Does the user provide only task-consistent facts and allow necessary clarification?
Checks C
Expected backend state and, for designated cases, final-response requirements
Does the evaluator accept the intended outcome and reject missing, wrong, or extra effects?
Appendix
Table 7: Artifacts in a final Thinkingbox-bench case and the corresponding construction or review question.
Case type
Required outcome
Incorrect effects rejected
Retail return or delivery exception
Correct order, return/refund/replacement, and support-ticket state under the applicable policy
Wrong order or item, ineligible refund, duplicate ticket, incorrect ticket status, or missing compensation record
Booking modification or cancellation
Correct booking dates, room/board attributes, charges or refund, and ticket or hotel escalation when required
Modification without availability or policy support, wrong fee, partial group update, or confidential partner information disclosed
Insurance billing, policy, or claim request
Correct policy-linked extension, driver/vehicle change, claim, document, cancellation, or reinstatement state
Identity or eligibility bypass, wrong effective date, unintended coverage change, or incorrect ticket resolution
Internal access, hardware, or software request
Correct employee, approval, access, asset, procurement, notification, and ticket records
Excess privilege, bypassed approval, wrong assignee or device, duplicate request, or incomplete multi-system update
Consulting operations request
Correct engagement-linked access, expense, onboarding, training, hardware, or travel outcome
Table 8: Representative evaluator targets drawn from the final task families. These are outcome conditions, not prescribed action sequences.
Parameter
Agent
Simulated User
Response Judge
temperature
1.0
0.3
0.0
max_completion_tokens
4096
4096
128
is_reasoning
True
False
False
reasoning_effort
medium
none
none
top- p
1.0
1.0
1.0
frequency penalty
0.0
0.0
0.0
Appendix
Table 9: Azure API inference parameters for the agent, fixed GPT-5.4-mini simulated user, and fixed GPT-5.4-mini response judge.
Parameter
Value
port
8000
data parallel size
8
tensor parallel size
1
maximum model length
65,536
reasoning parser
qwen3
automatic tool choice
True
Appendix
Table 10: vLLM serving parameters for the locally hosted Qwen-series models.
Qwen3.8-27B
Qwen3.6-27B
Qwen3.5-9B
Training paradigm
full-parameter
LoRA r16/α32
LoRA r16/α32
Training pool
187
168
157
Best checkpoint
50
20
26
Scheduled groups × rollouts
18×8
12×8
12×8
Parallelism
world24 / SP4 / DP6
world12 / SP2 / DP6
world12 / SP2 / DP6
GPUs (train + rollouts)
24 (shared)
12 + 4
12 + 4
Appendix
Table 11: Agent RL configurations. All three runs optimize the executable Thinkingbox verdict with token-mean GRPO. Group counts are scheduled; timeout-censored LoRA groups are dropped without replacement.
Decision source
Human A
Human B
Adjudicated reference
GPT-5.6 Sol, independent
92/120 (76.7%)
90/120 (75.0%)
92/120 (76.7%)
Claude Opus 5, independent
98/120 (81.7%)
94/120 (78.3%)
96/120 (80.0%)
Final pipeline label
92/120 (76.7%)
90/120 (75.0%)
94/120 (78.3%)
Appendix
Table 12: Agreement with human judgments on the 120-case review sample.
Automatic label
Human A
Human B
Ungrounded
32/60 (53.3%) [40.9, 65.4]
30/60 (50.0%) [37.7, 62.3]
Grounded
60/60 (100%) [94.0, 100.0]
60/60 (100%) [94.0, 100.0]
Appendix
Table 13: Human confirmation by final automatic label. Intervals are approximate 95% Wilson intervals.
Assistant/run
Turns
Ungrounded (%)
Grounded (%)
Disputed (%)
Claude Opus 5
527
1.71
98.10
0.19
GPT-5.4
512
10.35
89.26
0.39
GPT-5.6-sol
594
6.73
92.76
0.51
Claude Sonnet 4.6
1,152
6.77
93.06
0.17
Claude Opus 4.6
1,481
5.74
94.13
0.14
o3-pro
1,689
7.10
92.07
0.83
Appendix
Table 14: Automatic grounding labels by archived assistant/run group and domain. Percentages include Disputed cases in the denominator.
Figure 3: Automatic Ungrounded rates by archived run and domain. Cells show percentages and Ungrounded/all-turn counts, including Disputed turns in denominators.
Model
pass@1 (%)
pass^20 (%)
pass@20 (%)
0/20 tasks
20/20 tasks
Claude Opus 5
66.50 [62.78, 70.20]
47.53 [43.21, 51.84]
79.09 [75.37, 82.73]
106
241
GPT-5.4
65.36 [62.22, 68.50]
30.62 [26.71,34.52]
91.12 [88.45, 93.78]
45
128
GPT-5.6-sol
61.91 [58.65, 65.16]
22.00 [18.51, 25.49]
86.79 [83.70, 89.88]
67
82
Claude Sonnet 4.6
59.19 [55.23, 61.66]
25.36 [21.66, 29.06]
88.56 [84.93, 91.01]
58
102
GPT-6 Astra
58.31 [54.34, 62.28]
46.89 [42.56, 51.22]
71.01 [66.97, 75.05]
147
231
Kimi-K3
57.37 [54.31, 60.43]
17.60 [14.37, 20.83]
93.89 [91.52, 96.26]
31
68
Appendix
Table 15: Repeated-trial discovery and reliability on Thinkingbox-bench . The window next to the metrics denotes a 95% task-cluster bootstrap confidence interval. pass^20 uses the plug-in estimator in Equation 9 ; pass@20 requires at least one observed success. 0/20 and 20/20 means the number of tasks that have never been passed, and has always passed during all 20 attempts.
Figure 4: Pass@k progression of models used for evaluation
Figure 5: pass^k progression of models used for evaluation
Signal among executable-check failures
Trials
Share of failures
Failed trials accepted by weak observable evaluators
Clean termination
67,763
84.86%
Clean termination + state-changing tool call
64,586
80.88%
Above + no explicit error in final tool response
53,697
67.24%
Evidence reported by executable state/side-effect checks (ours)
Database hash mismatch
79,015
98.95%
Appendix
Table 16: Retrospective evaluator ablation study. The first block counts failed trials that nevertheless appear complete under weaker response- or action-level evaluators. The second block reports evidence exposed by our executable database comparison.
Figure 6: Domain failure-rate heatmap (%) on Thinkingbox-bench . Each cell is 100−pass@1 within the domain; darker red indicates higher failure rate.
Domain
Tool Usage
No State-Changing Action
Incomplete User Resolution
Wrong State Update
Retail
73.8
4.8
11.7
9.8
Travel
84.0
1.1
11.3
3.5
Auto insurance
51.3
3.8
15.1
29.9
Neobank internal IT
81.7
5.4
7.3
5.5
Consulting IT/HR
75.2
5.0
9.1
10.5
Appendix
Table 17: Failure distribution (%) by domain on Thinkingbox-bench , aggregated over the 15 models with available failure classifications. Each row reports percentages over failed trials in that domain.
Model
Avg. msgs. / trial
Avg. tool calls / trial
Avg. write calls / trial
Avg. tool errors / trial
Claude Opus 5
27.03
10.38
3.59
1.68
Claude Opus 4.6
31.20
9.99
3.21
1.57
Claude Sonnet 4.6
34.41
10.16
3.62
1.68
GPT-6 Astra
31.10
11.47
3.77
2.10
GPT-5.6-sol
33.31
10.55
3.70
1.72
GPT-5.4
29.80
11.05
3.91
1.98
Appendix
Table 18: Average trajectory-level interaction counts on Thinkingbox-bench . All columns are mean counts per trial, computed from parsed traces. Write calls are tool calls whose names indicate state-changing actions such as create, update, cancel, refund, or modify.
Model
Retail
Travel
Auto insurance
Neobank internal IT
Consulting IT/HR
Claude Opus 5
25.4k × 11.3
41.4k × 12.2
27.8k × 13.1
36.7k × 9.9
33.5k × 13.2
Claude Opus 4.6
14.9k × 8.7
27.3k × 7.6
15.8k × 10.1
25.1k × 7.0
17.1k × 7.6
Claude Sonnet 4.6
14.4k × 8.3
29.8k × 7.3
16.7k × 8.9
25.2k × 7.5
16.9k × 7.7
GPT-6 Astra
12.2k × 12.6
22.5k × 14.6
13.7k × 14.3
18.5k × 10.9
16.1k × 14.8
GPT-5.6-sol
12.6k × 11.2
26.7k × 13.6
13.9k × 13.0
20.9k × 9.1
16.8k × 13.8
GPT-5.4
14.2k × 8.1
27.4k × 7.2
14.6k × 9.6
21.3k × 6.7
18.2k × 7.4
Appendix
Table 19: Average token usage per model turn and average model turns per trial by model and domain. Each cell reports “Tokens/Turn x Turns”, with tokens in thousands; both averages use trials with token metadata. Tokens include input context and output tokens for each model invocation, with cached input counted once.
Model snapshot
Published
Direct prompt
Thinkingbox [95% CI]
GPT-4o (Aug 2024)
87.2
85.37
84.76 [79.27, 90.24]
GPT-4o-mini (July 2024)
83.5
81.10
85.37 [79.88, 90.24]
GPT-4-Turbo (April 2024)
86.6
87.20
83.54 [77.44, 89.02]
GPT-3.5-Turbo (Nov 2023)
70.7
65.85
69.51 [62.20, 76.22]
Appendix
Table 20: HumanEval+ pass@1 (%) for dated OpenAI snapshots. Published values are from the September 16, 2026 EvalPlus leaderboard snapshot. Direct prompt and Thinkingbox use one attempt per task; brackets give 95% task-bootstrap intervals.
Thinkingbox-bench
HumanEval+, no interpreter
HumanEval+, interpreter
Model
pass@1
pass@1 [95% CI]
All-five
pass@1 [95% CI]
All-five
Claude Opus 5
66.50
95.24 [91.95, 98.17]
94.51
94.51 [90.85, 97.56]
93.29
GPT-5.4
65.36
93.90 [90.61, 96.71]
88.41
94.51 [91.22, 97.44]
91.46
GPT-5.6-sol
61.91
93.29 [89.27, 96.83]
92.68
93.90 [90.24, 97.07]
92.07
Claude Sonnet 4.6
59.19
92.56 [88.78, 95.98]
89.02
91.59 [87.32, 95.37]
89.02
GPT-6 Astra
58.31
94.76 [91.10, 97.80]
94.51
95.12 [91.59, 98.17]
94.51
Appendix
Table 21: Pass@1 and all-five success (%). Thinkingbox-bench pass@1 comes from Table 15 ( N=20 ); HumanEval+ uses five attempts per task. Brackets give 95% task-cluster intervals. Claude request settings and returned model IDs are unverified; † additionally marks Opus 4.6’s unusually repetitive outputs.
Model
Difference
95% CI
Attempts using interpreter
Claude Opus 5
−0.73
[−2.20,0.73]
820/820
GPT-5.4
+0.61
[−1.59,2.93]
820/820
GPT-5.6-sol
+0.61
[−0.61,2.07]
820/820
Claude Sonnet 4.6
−0.98
[−3.41,1.34]
820/820
GPT-6 Astra
+0.37
[−1.22,2.20]
820/820
GPT-5.2
+1.46
[−0.12,3.17]
662/820
Appendix
Table 22: Interpreter minus no-interpreter pass@1, in percentage points, with paired 95% task-cluster intervals. Usage counts attempts with at least one interpreter call, out of 820. Claude settings are unverified as noted in Table 21 .
Figure 7: Estimated cost per successful attempt against pass@1, colored by vendor. Points ringed in blue are the Pareto cost frontier, the dashed staircase traces the best pass@1 obtainable at or below each price and shades the region it dominates. Four models in Table 23 lie off-scale to the right.
#
Model
Rates: I / O / R / W
pass@1 (%)
Succ.
20/20
C^m ($)
$ / succ.
$ / all-20 task
1
GPT-5.6-sol
2.00 / 10.00 / 0.20 / 2.50
61.91
6,278
82
800.00
0.127
9.76
2
GPT-5.4
1.25 / 7.50 / 0.12 / –
65.36
6,628
128
869.80
0.131
6.80
3
Kimi-K2.6
0.57 / 2.40 / 0.11 / –
37.66
3,819
16
543.60
0.142
33.98
4
Qwen3.8-27B
0.20 / 2.50 / 0.05 / –
51.70
5,242
38
925.80
0.177
24.36
5
GPT-5.2
0.87 / 7.00 / 0.08 / –
46.28
4,693
44
878.00
0.187
19.95
6
Kimi-K3
2.10 / 10.95 / 0.21 / –
57.37
5,817
68
1,406.40
0.242
20.68
Appendix
Table 23: Campaign-cost estimates for eighteen models, as per Equation 14 , sorted by cost per successful attempt.
Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover complex task settings, they still largely assume clean, stable, and trustworthy tool environments, leaving tool-environment unreliability insufficiently examined. We introduce ToolBench-X, a benchmark for evaluating agents under recoverable reliability hazards. ToolBench-X contains executable multi-step tasks across diverse domains and sequential, parallel, and mixed workflows, each paired with deterministic tools and a canonical final answer for automatic evaluation. Starting from clean tool environments, ToolBench-X injects five structured hazard types: Specification Drift, Invocation Error, Execution Failure, Output Drift, and Cross-source Conflict. Crucially, each injected instance remains solvable through at least one valid recovery path, such as retrying, fallback, verification, or cross-checking. Experiments reveal a substantial reliability gap: agents that perform well with reliable tools often fail under recoverable hazards. Further analysis shows that failures are driven less by tool-use volume or inference budget than by limited hazard diagnosis and ineffective recovery. Targeted recovery hints recover many failed tasks, while test-time scaling yields more limited gains. These results suggest that tool-use evaluation should move beyond function-call accuracy toward task completion under unreliable tool environments. The code and data is available at https://github.com/Foreverskyou/ToolBench-X.
Yang Tian, Zhengpeng Shi, Yu Zhou +1
School of AI, Shanghai Jiao Tong University1 · School of Cyber Science and Engineering, Southeast University2
Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static prompt evaluation misses failures that only appear when agents operate over persistent state. Existing interactive benchmarks have advanced agent evaluation significantly, but most initialize tasks from clean state and do not systematically test how agents handle pre-existing partial, stale, or conflicting artifacts. We present \textbf{ClawForge}, a generator-backed benchmark framework for executable command-line workflows under state conflict. The framework compiles scenario templates, grounded slots, initialized state, reference trajectories, and validators into reproducible task specifications, and evaluates agents step by step over persistent workflow surfaces using normalized end state and observable side effects rather than exact trajectory matching. We instantiate this framework as the ClawForge-Bench (17 scenarios, 6 ability categories). Results across seven frontier models show that the best model reaches only 45.3% strict accuracy, wrong-state replacement remains below 17% for all models, and the widest model separation (17% to 90%) is driven by whether agents inspect existing state before acting. Partial-credit and step-efficiency analyses further reveal that many failures are near-miss closures rather than early breakdowns, and that models exhibit qualitatively different failure styles under state conflict.
Yuxiang Lai, Peng Xia, Haonian Ji +8
University of North Carolina at Chapel Hill · Stanford University · University of Southern California +2
Tool-using agents increasingly rely on external tools to complete multi-step tasks, but tool returns can fail in different ways and require different recovery actions. Existing robustness studies often use uncertainty-based measures to detect when an agent becomes unreliable. These measures can reveal that something has gone wrong, but they do not directly identify the type of tool failure or the appropriate response. We address this limitation by analyzing tool failures at the moment a return enters the agent context. Our approach combines two complementary signals. The first compares the likelihood of the returned content under the tool schema and under the full trajectory prefix. The second measures the agent's probability distribution over its legal next actions. We evaluate the approach by injecting incomplete and inconsistent returns into a retail customer-service benchmark. The results show that likelihood-based signals clearly capture incomplete returns and some direct inconsistencies, while action-based signals reveal how strongly a failure changes the next decision. Some failures that are weak under likelihood signals can still redirect the agent toward state-changing actions. These findings show that tool failures can be recognized at the return boundary, but reliable diagnosis requires combining multiple signals.
Jiachen Xu, Torben Bach Pedersen, Zhongming Yao +2
Department of Computer Science Aalborg University Aalborg, Denmark · College of Computer Science and Technology Zhejiang University Hangzhou, China · School of Information Science and Engineering Northeastern University Shenyang, China