cs.AISep 27, 2026

AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

Authors: Tianzhuo Yang, Zirui Mi, Yantao Huang, Guoxi Zhang, Jiawei Chen, Yaodong Yang, Jingwei Yi

Organizations: Peking University · Beijing Academy of Artificial Intelligence

Abstract

Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5% of routine-looking unauthorized actions yet completes only 28.7% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

    Jul 15, 2026Aryan Keluskar, Amrita Bhattacharjee, Huan LiuLarge Language Model SafetyAgentic Deployments

  2. Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

    Jul 31, 2026Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan +2RiskSchema

  3. Agent Safety Is Action Alignment

    Jun 27, 2026Shawn Li, Yue ZhaoRefusals