One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Organizations: University of Pittsburgh · Northwestern University · University of California, Irvine · Microsoft
Abstract
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across business scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Our experiments reveal that even the strongest proprietary and open-weight models show steep reliability drops: Claude Opus 5 falls from 66.50% pass@1 to 47.53% pass^20, and Kimi-K3 from 57.37% pass@1 to 17.60% pass^20. Moreover, many failed trials terminate cleanly after valid state-changing actions, so response- or tool-call-level signals poorly proxy end-to-end completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox (https://github.com/microsoft/thinkingbox) and Thinkingbox-bench (https://github.com/microsoft/thinkingbox-data).
Figures & tables
| Benchmark | Primary domain | Tools/APIs | User dialogue | Stateful backend | Side-effect checks | MCP servers |
|---|---|---|---|---|---|---|
| SWE-bench ( Jimenez et al., 2024 ) | Code repair | |||||
| BFCL ( Patil et al., 2025 ) | Function calling | |||||
| ToolBench / API-Bank ( Qin et al., 2024 ; Li et al., 2023 ) | API tool use | |||||
| WebArena / OSWorld ( Zhou et al., 2024 ; Xie et al., 2024 ) | Web/desktop control | |||||
| AppWorld ( Trivedi et al., 2024 ) | App APIs / coding agents | |||||
| MCP-Atlas ( Bandi et al., 2026 ) | Real MCP servers |
| Retail | Booking | Insurance | Neobank | Consulting | |
| Tasks | 98 | 104 | 100 | 104 | 101 |
| Backend systems | 11 | 8 | 7 | 3 | 18 |
| Databases (tables / rows) | 22 / 86 | 17 / 98 | 14 / 72 | 20 / 151 | 30 / 231 |
| Agent tools (write / read) | 16 / 17 | 10 / 28 | 14 / 19 | 13 / 19 | 13 / 14 |
| Policy (words) | 945 | 3,684 | 2,471 | 3,392 | 1,747 |
| Knowledge base (documents) | 9 | 11 | 8 | 8 | 9 |
| Domain | Scenario | Required agent behavior | Executable checks |
|---|---|---|---|
| Retail / e-commerce | User asks to change or refund part of an order. | Identify the correct order/item, verify eligibility, request missing confirmation, and update only the relevant record. | Correct order state; refund/order side effect; no unrelated customer or item modified. |
| Travel / hospitality | User requests a booking change under date, room, or policy constraints. | Check reservation, availability, and change policy before modifying booking or explaining denial. | Correct reservation state; price/fee side effect if applicable; policy-compliant dialogue. |
| Auto insurance | User reports or updates a claim. | Verify policy and vehicle/incident details, collect missing information, and create or update the claim. | Correct ticket/claim state; no coverage mutation unless allowed. |
| Neobank internal IT | Employee requests access to an internal application. | Verify employee role, existing access, and required approvals before provisioning or escalating. | Correct access, approval, and ticket state; no unauthorized privileges or unrelated records changed. |
| Consulting IT/HR | Employee requests enrollment in remaining onboarding courses. | Check existing enrollment, add missing courses, and keep the onboarding ticket pending until completion. | Correct course enrollment and ticket state; no duplicate enrollment or premature closure. |
| Model | Size | Retail (98) | Auto (100) | Booking (104) | Bank (104) | Consulting (101) | Average |
|---|---|---|---|---|---|---|---|
| Proprietary models | |||||||
| Claude Opus 5 | – | 80.71 3.68 | 65.80 4.24 | 49.95 4.38 | 70.62 4.10 | 66.19 4.32 | 66.50 1.91 |
| GPT-5.4 | – | 76.33 3.53 | 62.65 3.71 | 68.12 3.46 | 65.34 3.17 | 54.60 4.00 | 65.36 1.63 |
| GPT-5.6-sol | – | 67.65 3.55 | 65.30 3.66 | 60.34 3.65 | 59.09 3.64 | 57.52 4.02 | 61.91 1.66 |
| Claude Sonnet 4.6 | – | 72.35 3.65 | 54.40 4.03 | 58.94 3.8 | 56.39 3.40 | 54.31 3.73 | 59.19 1.69 |
| GPT-6 Astra | – | 71.73 4.38 | 46.55 4.31 | 55.87 4.68 | 60.87 4.41 | 56.83 4.56 | 58.31 2.03 |
| Model | Tool Usage | No State-Changing Action | Incomplete User Resolution | Wrong State Update |
|---|---|---|---|---|
| Claude Opus 5 | 96.4 | 0.8 | 0.8 | 2.0 |
| GPT-5.4 | 89.6 | 1.6 | 0.5 | 8.3 |
| GPT-5.6-sol | 86.6 | 8.6 | 0.2 | 4.6 |
| Claude Sonnet 4.6 | 84.0 | 2.8 | 4.2 | 8.9 |
| GPT-6 Astra | 97.0 | 0.8 | 0.3 | 1.9 |
| Claude Opus 4.6 | 78.1 | 3.0 | 10.2 | 8.7 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Backend only | Backend + rubric | Task families represented in the final set |
|---|---|---|---|
| Retail / e-commerce | 98 | 0 | Delivery delay and exception handling; delivered-but-missing and return-to-sender cases; returns, refunds, exchanges, and warranty claims; installation scheduling and cancellation; order cancellation, promotions, and membership changes. |
| Travel / hospitality | 89 | 15 | Individual, corporate, and group booking changes; payment recovery; cancellations and refunds; corporate invoices and account benefits; group billing and services; hotel-partner verification and discrepancies; special requests and post-stay complaints. |
| Auto insurance | 100 | 0 | Billing extensions and arrangements; proof-of-insurance documents; adding, removing, or updating drivers and vehicles; first notice of loss and claim intake; listed-driver requests; reinstatement and policy cancellation. |
| Neobank internal IT support | 89 | 15 | Employee access and approval requests; password and account-security actions; production-incident access; hardware troubleshooting, assignment, replacement, and procurement; software and license requests; policy-information questions. |
| Consulting IT / HR support | 101 | 0 | Client-system and document access; software provisioning; expenses; hardware requests; employee onboarding; training enrollment; engagement and approval checks; corporate-travel policy and escalation. |
| Total | 477 | 30 | 507 executable cases in total. |
| Artifact | Contents | Review question |
|---|---|---|
| User goal | Initial request plus facts the simulated user can provide during follow-up | Is the request natural, internally consistent, and resolvable without access to hidden evaluator information? |
| Initial state | Synthetic records in the domain backend, including existing tickets and related business objects | Do all referenced identifiers resolve, and do cross-system records agree before the agent acts? |
| Policy context | Domain operating manual and the fixed evaluation time | Does the policy determine eligibility, approvals, disclosures, and allowed actions without exposing the golden result? |
| Tools | MCP-compatible read and write operations over the isolated domain services | Can the required evidence be retrieved and the intended outcome be executed using available tools? |
| User simulation | Task-specific user role used for on-policy follow-up dialogue | Does the user provide only task-consistent facts and allow necessary clarification? |
| Checks | Expected backend state and, for designated cases, final-response requirements | Does the evaluator accept the intended outcome and reject missing, wrong, or extra effects? |
| Case type | Required outcome | Incorrect effects rejected |
|---|---|---|
| Retail return or delivery exception | Correct order, return/refund/replacement, and support-ticket state under the applicable policy | Wrong order or item, ineligible refund, duplicate ticket, incorrect ticket status, or missing compensation record |
| Booking modification or cancellation | Correct booking dates, room/board attributes, charges or refund, and ticket or hotel escalation when required | Modification without availability or policy support, wrong fee, partial group update, or confidential partner information disclosed |
| Insurance billing, policy, or claim request | Correct policy-linked extension, driver/vehicle change, claim, document, cancellation, or reinstatement state | Identity or eligibility bypass, wrong effective date, unintended coverage change, or incorrect ticket resolution |
| Internal access, hardware, or software request | Correct employee, approval, access, asset, procurement, notification, and ticket records | Excess privilege, bypassed approval, wrong assignee or device, duplicate request, or incomplete multi-system update |
| Consulting operations request | Correct engagement-linked access, expense, onboarding, training, hardware, or travel outcome | Missing prerequisite, incorrect approval path, inconsistent cross-system records, or premature ticket closure |
| Parameter | Agent | Simulated User | Response Judge |
|---|---|---|---|
| temperature | 1.0 | 0.3 | 0.0 |
| max_completion_tokens | 4096 | 4096 | 128 |
| is_reasoning | True | False | False |
| reasoning_effort | medium | none | none |
| top- | 1.0 | 1.0 | 1.0 |
| frequency penalty | 0.0 | 0.0 | 0.0 |
| Parameter | Value |
|---|---|
| port | 8000 |
| data parallel size | 8 |
| tensor parallel size | 1 |
| maximum model length | 65,536 |
| reasoning parser | qwen3 |
| automatic tool choice | True |
| Qwen3.8-27B | Qwen3.6-27B | Qwen3.5-9B | |
| Training paradigm | full-parameter | LoRA | LoRA |
| Training pool | 187 | 168 | 157 |
| Best checkpoint | 50 | 20 | 26 |
| Scheduled groups rollouts | |||
| Parallelism | world24 / SP4 / DP6 | world12 / SP2 / DP6 | world12 / SP2 / DP6 |
| GPUs (train + rollouts) | 24 (shared) | 12 + 4 | 12 + 4 |
| Decision source | Human A | Human B | Adjudicated reference |
|---|---|---|---|
| GPT-5.6 Sol, independent | 92/120 (76.7%) | 90/120 (75.0%) | 92/120 (76.7%) |
| Claude Opus 5, independent | 98/120 (81.7%) | 94/120 (78.3%) | 96/120 (80.0%) |
| Final pipeline label | 92/120 (76.7%) | 90/120 (75.0%) | 94/120 (78.3%) |
| Automatic label | Human A | Human B |
|---|---|---|
| Ungrounded | 32/60 (53.3%) [40.9, 65.4] | 30/60 (50.0%) [37.7, 62.3] |
| Grounded | 60/60 (100%) [94.0, 100.0] | 60/60 (100%) [94.0, 100.0] |
| Assistant/run | Turns | Ungrounded (%) | Grounded (%) | Disputed (%) |
|---|---|---|---|---|
| Claude Opus 5 | 527 | 1.71 | 98.10 | 0.19 |
| GPT-5.4 | 512 | 10.35 | 89.26 | 0.39 |
| GPT-5.6-sol | 594 | 6.73 | 92.76 | 0.51 |
| Claude Sonnet 4.6 | 1,152 | 6.77 | 93.06 | 0.17 |
| Claude Opus 4.6 | 1,481 | 5.74 | 94.13 | 0.14 |
| o3-pro | 1,689 | 7.10 | 92.07 | 0.83 |
| Model | pass@1 (%) | (%) | pass@20 (%) | 0/20 tasks | 20/20 tasks |
|---|---|---|---|---|---|
| Claude Opus 5 | 66.50 [62.78, 70.20] | 47.53 [43.21, 51.84] | 79.09 [75.37, 82.73] | 106 | 241 |
| GPT-5.4 | 65.36 [62.22, 68.50] | 30.62 [26.71,34.52] | 91.12 [88.45, 93.78] | 45 | 128 |
| GPT-5.6-sol | 61.91 [58.65, 65.16] | 22.00 [18.51, 25.49] | 86.79 [83.70, 89.88] | 67 | 82 |
| Claude Sonnet 4.6 | 59.19 [55.23, 61.66] | 25.36 [21.66, 29.06] | 88.56 [84.93, 91.01] | 58 | 102 |
| GPT-6 Astra | 58.31 [54.34, 62.28] | 46.89 [42.56, 51.22] | 71.01 [66.97, 75.05] | 147 | 231 |
| Kimi-K3 | 57.37 [54.31, 60.43] | 17.60 [14.37, 20.83] | 93.89 [91.52, 96.26] | 31 | 68 |
| Signal among executable-check failures | Trials | Share of failures |
|---|---|---|
| Failed trials accepted by weak observable evaluators | ||
| Clean termination | 67,763 | 84.86% |
| Clean termination + state-changing tool call | 64,586 | 80.88% |
| Above + no explicit error in final tool response | 53,697 | 67.24% |
| Evidence reported by executable state/side-effect checks (ours) | ||
| Database hash mismatch | 79,015 | 98.95% |
| Domain | Tool Usage | No State-Changing Action | Incomplete User Resolution | Wrong State Update |
|---|---|---|---|---|
| Retail | 73.8 | 4.8 | 11.7 | 9.8 |
| Travel | 84.0 | 1.1 | 11.3 | 3.5 |
| Auto insurance | 51.3 | 3.8 | 15.1 | 29.9 |
| Neobank internal IT | 81.7 | 5.4 | 7.3 | 5.5 |
| Consulting IT/HR | 75.2 | 5.0 | 9.1 | 10.5 |
| Model | Avg. msgs. / trial | Avg. tool calls / trial | Avg. write calls / trial | Avg. tool errors / trial |
|---|---|---|---|---|
| Claude Opus 5 | 27.03 | 10.38 | 3.59 | 1.68 |
| Claude Opus 4.6 | 31.20 | 9.99 | 3.21 | 1.57 |
| Claude Sonnet 4.6 | 34.41 | 10.16 | 3.62 | 1.68 |
| GPT-6 Astra | 31.10 | 11.47 | 3.77 | 2.10 |
| GPT-5.6-sol | 33.31 | 10.55 | 3.70 | 1.72 |
| GPT-5.4 | 29.80 | 11.05 | 3.91 | 1.98 |
| Model | Retail | Travel | Auto insurance | Neobank internal IT | Consulting IT/HR |
|---|---|---|---|---|---|
| Claude Opus 5 | 25.4k 11.3 | 41.4k 12.2 | 27.8k 13.1 | 36.7k 9.9 | 33.5k 13.2 |
| Claude Opus 4.6 | 14.9k 8.7 | 27.3k 7.6 | 15.8k 10.1 | 25.1k 7.0 | 17.1k 7.6 |
| Claude Sonnet 4.6 | 14.4k 8.3 | 29.8k 7.3 | 16.7k 8.9 | 25.2k 7.5 | 16.9k 7.7 |
| GPT-6 Astra | 12.2k 12.6 | 22.5k 14.6 | 13.7k 14.3 | 18.5k 10.9 | 16.1k 14.8 |
| GPT-5.6-sol | 12.6k 11.2 | 26.7k 13.6 | 13.9k 13.0 | 20.9k 9.1 | 16.8k 13.8 |
| GPT-5.4 | 14.2k 8.1 | 27.4k 7.2 | 14.6k 9.6 | 21.3k 6.7 | 18.2k 7.4 |
| Model snapshot | Published | Direct prompt | Thinkingbox [95% CI] |
|---|---|---|---|
| GPT-4o (Aug 2024) | 87.2 | 85.37 | 84.76 [79.27, 90.24] |
| GPT-4o-mini (July 2024) | 83.5 | 81.10 | 85.37 [79.88, 90.24] |
| GPT-4-Turbo (April 2024) | 86.6 | 87.20 | 83.54 [77.44, 89.02] |
| GPT-3.5-Turbo (Nov 2023) | 70.7 | 65.85 | 69.51 [62.20, 76.22] |
| Thinkingbox-bench | HumanEval+, no interpreter | HumanEval+, interpreter | |||
|---|---|---|---|---|---|
| Model | pass@1 | pass@1 [95% CI] | All-five | pass@1 [95% CI] | All-five |
| Claude Opus 5 | 66.50 | 95.24 [91.95, 98.17] | 94.51 | 94.51 [90.85, 97.56] | 93.29 |
| GPT-5.4 | 65.36 | 93.90 [90.61, 96.71] | 88.41 | 94.51 [91.22, 97.44] | 91.46 |
| GPT-5.6-sol | 61.91 | 93.29 [89.27, 96.83] | 92.68 | 93.90 [90.24, 97.07] | 92.07 |
| Claude Sonnet 4.6 | 59.19 | 92.56 [88.78, 95.98] | 89.02 | 91.59 [87.32, 95.37] | 89.02 |
| GPT-6 Astra | 58.31 | 94.76 [91.10, 97.80] | 94.51 | 95.12 [91.59, 98.17] | 94.51 |
| Model | Difference | 95% CI | Attempts using interpreter |
|---|---|---|---|
| Claude Opus 5 | 820/820 | ||
| GPT-5.4 | 820/820 | ||
| GPT-5.6-sol | 820/820 | ||
| Claude Sonnet 4.6 | 820/820 | ||
| GPT-6 Astra | 820/820 | ||
| GPT-5.2 | 662/820 |
| # | Model | Rates: I / O / R / W | pass@1 (%) | Succ. | ($) | $ / succ. | $ / all-20 task | |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6-sol | 2.00 / 10.00 / 0.20 / 2.50 | 61.91 | 6,278 | 82 | 800.00 | 0.127 | 9.76 |
| 2 | GPT-5.4 | 1.25 / 7.50 / 0.12 / – | 65.36 | 6,628 | 128 | 869.80 | 0.131 | 6.80 |
| 3 | Kimi-K2.6 | 0.57 / 2.40 / 0.11 / – | 37.66 | 3,819 | 16 | 543.60 | 0.142 | 33.98 |
| 4 | Qwen3.8-27B | 0.20 / 2.50 / 0.05 / – | 51.70 | 5,242 | 38 | 925.80 | 0.177 | 24.36 |
| 5 | GPT-5.2 | 0.87 / 7.00 / 0.08 / – | 46.28 | 4,693 | 44 | 878.00 | 0.187 | 19.95 |
| 6 | Kimi-K3 | 2.10 / 10.95 / 0.21 / – | 57.37 | 5,817 | 68 | 1,406.40 | 0.242 | 20.68 |