cs.LGSep 28, 2026

PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

Authors: Ding Jia, Wei Liu, Xianglong Du, Yingjie Li, Yingqing Yang, Huili Yu, Zhangsong Zhan, Chu Zhou

Organizations: State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile · Independent Researcher

Abstract

The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

    Aug 6, 2026Wenhao Lin, Chenyu Yu, Xingwei Lin +6Runtime Safety FilteringLarge Language Model Agents

  2. SafeAgent: A Runtime Protection Architecture for Agentic Systems

    Apr 19, 2026Hailin Liu, Eugene Ilyushin, Jie Ni +1Runtime Safety FilteringLarge Language Model Safety

  3. PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

    Sep 23, 2026Jiapeng Sun, Yujin Zhou, Han Zhu +4ProactiveHazard