cs.AIOct 6, 2026

POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

Authors: Yunju Kang, Seonghyeon Cho, Irene Li, Yeo-Chan Yoon, Chanjun Park

Organizations: Soongsil University · Tokyo University · Jeju National University

Abstract

LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on τ2τ^2-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

    Sep 27, 2026Tianzhuo Yang, Zirui Mi, Yantao Huang +4RefusalsEvaluation Benchmarks

  2. Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

    Jul 31, 2026Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan +2RiskSchema

  3. DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

    Aug 6, 2026Wenhao Lin, Chenyu Yu, Xingwei Lin +6Runtime Safety FilteringLarge Language Model Agents