cs.CLAug 20, 2026

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Authors: Zhuochun Li, Youngmin Ko, Ali Keramati, Tuhin Kundu, Liang-Chun Tsai, Nicola Ferri, Mirco Milletari, Jiaxiang Liu, +6 more

Organizations: University of Pittsburgh · Northwestern University · University of California, Irvine · Microsoft

Abstract

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across business scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Our experiments reveal that even the strongest proprietary and open-weight models show steep reliability drops: Claude Opus 5 falls from 66.50% pass@1 to 47.53% pass^20, and Kimi-K3 from 57.37% pass@1 to 17.60% pass^20. Moreover, many failed trials terminate cleanly after valid state-changing actions, so response- or tool-call-level signals poorly proxy end-to-end completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox (https://github.com/microsoft/thinkingbox) and Thinkingbox-bench (https://github.com/microsoft/thinkingbox-data).

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

    Jun 24, 2026Yang Tian, Zhengpeng Shi, Yu Zhou +1Controlling Tool UseAgentic Benchmarks

  2. ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

    May 13, 2026Yuxiang Lai, Peng Xia, Haonian Ji +8Agentic BenchmarksAgentic Workflow Design

  3. Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

    Aug 24, 2026Jiachen Xu, Torben Bach Pedersen, Zhongming Yao +2Controlling Tool UseRobustness Verification