The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.
Figures & tables
Figure 1: The PROACT-Agent ecosystem. (a) PROACT-Agent formulates risk as a non-Markovian prefix-evaluation task, allowing intervention on observed context before the next LLM inference. (b) The framework combines trajectory context, prefix-level evaluation, and monotonic causal rectification. (c) PROACT-Bench contains 155,780 bilingual states with available or inferred auxiliary rationales and model-adjudicated safety labels.
Figure 2: Overview of the PROACT-Agent framework, which transforms static logs into a causally-rigorous streaming dataset across three stages: (1) Progressive Trajectory Unrolling (PTU) unrolls trajectories into non-Markovian evaluation states; (2) Reasoning-Augmented Causal Rectification (RACR) infers missing auxiliary rationales and rectifies labels via Monotonic Causal Consistency (MCC); and (3) Culturally-Aware Data Localization (CADL) maps states across English and Chinese contexts with Tripartite Rigid Verification (TRV) checks.
Figure 3: Overview of the PROACT-Bench dataset. (a) Scale and language: 155,780 prefix-level states across English and localized Chinese environments. (b) Quality verification: Data Quality Score (DQS) summarizes model-committee agreement and contested labels within the adjudication pipeline. (c) Taxonomic coverage: the benchmark maps six safety risk categories across a range of operational environments.
Method
Overall
English Subset
Chinese Subset
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Closed-source LLMs
Qwen3-Max [ 1 ]
87.52
80.25
98.25
88.34
87.04
78.66
98.62
87.52
88.00
81.80
97.90
89.13
Gemini 3 Flash [ 8 ]
88.73
84.30
94.22
88.98
88.08
83.15
92.98
87.79
89.38
85.35
95.34
90.07
Open-source LLMs
DeepSeek-V3.2 [ 19 ]
86.24
79.77
95.74
87.03
81.55
73.12
94.84
82.57
90.96
86.93
96.58
91.50
Table 1: Broad Mixed-Source Comparison. Accuracy (Acc.), unsafe Precision (Pre.), Recall (Rec.), and F1 (%) on the Overall, English, and Chinese subsets of the mixed-source comparison protocol. Best results are bold ; second-best results are underlined .
Strict Root-Grouped
Complete Source-Held-Out
Method
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Qwen3Guard-Gen-8B
94.67
88.59
93.96
91.20
75.09
76.67
66.77
71.38
TS-Guard
90.36
79.40
90.75
84.70
75.06
74.81
69.94
72.29
PROACT-Agent
97.85
96.44
96.23
96.34
92.37
95.52
87.72
91.46
Table 2: Strict Generalization. Acc., unsafe Pre., Rec., and F1 (%) on 31,156 root-grouped and 3,396 complete source-held-out evaluation states. Invalid outputs count as errors.
Strict Root-Grouped
Complete Source-Held-Out
Method
First-Unsafe
Exact
First-Unsafe
Exact
Qwen3Guard-Gen-8B
94.39
93.18
62.93
57.99
TS-Guard
91.05
90.14
64.67
60.42
PROACT-Agent
97.38
97.12
91.84
90.63
Table 3: Temporal Boundary Detection (%). First-unsafe recall requires detecting the gold first-unsafe prefix as unsafe; exact-boundary detection additionally requires every earlier safe prefix to remain unblocked. The denominator is root-language trajectories with an unsafe boundary, with EN/ZH evaluated separately. Gold boundaries are model-adjudicated; invalid predictions count as errors.
Defense
Successful Attacks ↓
Banking ASR ↓
Slack ASR ↓
Travel ASR ↓
Workspace ASR ↓
Overall ASR ↓
No guard
2,173 / 10,439
36.87%
56.62%
40.06%
5.16%
20.82%
+ PROACT-Agent
42 / 10,439
0.00%
0.00%
2.73%
0.00%
0.40%
Table 4: AgentDojo Closed-Loop Results on Non-DoS Attacks. Suite-wise and overall targeted attack success rates (ASR) across four suites. Each paired logical case is identified by suite, attack, user task, and injection task; the overall evaluation contains 10,439 cases.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Fixed ID 20%
Complete OOD
Training
Selected Step
Macro-F1 ↑
FBR ↓
FAR ↓
Macro-F1 ↑
FBR ↓
FAR ↓
60%
8,500
98.08
0.96
3.10
90.55
3.69
15.76
80%
8,000
98.24
0.91
2.79
91.78
4.13
12.72
Δ (80 − 60)
—
+0.16
−0.05
−0.31
+1.23
+0.44
−3.04
Appendix
Table 6: Matched Two-Source Training-Size Sensitivity. Selected checkpoints from AgentAlign+ToolSafety training runs share fixed-ID (30,470 states) and complete-OOD (3,396 states) evaluations. Metrics are percentages; Δ denotes 80% minus 60% in percentage points. Invalid predictions are zero in all settings.
Backbone
Acc.
Pre.
Rec.
F1
Invalid
Qwen3-1.7B
96.14
95.30
96.73
96.01
0
Qwen3-4B
96.93
96.51
97.11
96.81
0
Qwen3-8B
96.91
96.77
96.80
96.78
4
Appendix
Table 7: Qwen3 Backbone-Scale Robustness. Accuracy, unsafe Precision, Recall, and F1 (%) on the same mixed-source evaluation set for selected Qwen3 checkpoints. Invalid predictions count as errors.
Backbone
Acc.
Pre.
Rec.
F1
Invalid
Qwen2.5-7B-Instruct
97.11
96.30
97.74
97.01
0
Qwen3-8B
96.91
96.77
96.80
96.78
4
Qwen3.5-9B
97.16
96.85
97.25
97.05
0
Appendix
Table 8: Cross-Generation Backbone Robustness. Accuracy, unsafe Precision, Recall, and F1 (%) on the same mixed-source evaluation set. Model generation and parameter count both vary, so this is not a controlled scaling experiment. Qwen3-8B intentionally bridges this comparison and Table 7 .
PROACT backbone
Standard P50 / P95
Memory (GiB)
Slope / 1K (ms)
4K+ P50 / P95
6K+ P50 / P95
Qwen3-1.7B
403 / 889
4.69
20.8
575 / 1,138
657 / 1,151
Qwen3-4B
543 / 1,137
8.78
62.1
894 / 1,714
1,118 / 1,723
Qwen2.5-7B-Instruct
474 / 976
15.63
86.6
899.8 / 1,629
1,159 / 1,633
Appendix
Table 9: Guard-Only Runtime Efficiency and Long-Context Scaling. Measurements use one A800-SXM4-80GB at batch size 1. Standard latency summarizes the overall guard-only workload; 4K+ and 6K+ summarize long-context subsets with input lengths exceeding the corresponding token thresholds. Slope is the empirical latency-growth trend per additional 1K input tokens. Latency is in milliseconds; memory is P95 peak allocated memory in GiB.
Method
Overall
English Subset
Chinese Subset
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
W/o CADL
91.19
85.89
97.79
91.45
90.21
83.23
98.62
90.27
92.18
88.54
97.02
92.59
W/o RACR
92.09
91.65
91.98
91.82
91.57
87.83
94.84
91.20
92.62
95.72
89.35
92.42
W/o PTU
93.48
93.67
92.74
93.20
92.52
90.12
94.09
92.06
94.44
97.30
91.50
94.31
PROACT-Agent (Full)
94.53
93.16
95.68
94.40
93.57
89.56
97.39
93.31
95.50
96.87
94.10
95.46
Appendix
Table 10: Mixed-Source Ablation. Impact of independently removing the three core operators (CADL, RACR, and PTU) from the PROACT-Agent framework under the mixed-source comparison protocol. All variants are trained using the Qwen2.5-7B [ 35 ] backbone. The full pipeline achieves the best Accuracy and F1-score across all evaluation cohorts, while maintaining a strong Precision–Recall balance. Unsafe is treated as the positive class. Best results are shown in bold and second-best results are underlined .
Figure 4: The system prompt deployed to the cognitive parser ( Mparse ), which is used in the PTU module.
Figure 5: The system prompt for reasoning recovery used in the RACR module.
Figure 6: The system prompt for the diverse evaluator committee ( M∈M ) used in the RACR module - Part 1: Persona definition and baseline safety rules.
Figure 7: The system prompt for the diverse evaluator committee ( M∈M ) used in the RACR module - Part 2: High-risk categorization and output formatting constraints.
Figure 8: The system prompt for the information-augmented arbitrator ( J ) used in the RACR module.
Figure 9: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 1: Persona definition, core principles, and detailed mapping rules.
Figure 10: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 2: An abridged few-shot English baseline example reproduced from the localization template.
Figure 11: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 3: Explicit rationale formulation demonstrating how entities and risks are mapped onto the Chinese cultural manifold.
Figure 12: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 4: An abridged localized JSON few-shot example reproduced from the localization template.
Figure 13: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 5: Dynamic data injection and the constrained two-step output workflow. Template placeholders are rendered before model invocation.
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many remain reactive: they primarily assess the apparent safety of the current action, lacking an explicit model of how risk evolves across the trajectory. This limitation creates a critical blind spot for long-horizon risks, where individually benign-looking actions can gradually drift the agent toward hazardous states. In response, we propose DreamGuard, a proactive guardrail for LLM agents built around a risk-aware world model. The world model maintains a compact recurrent latent state over the trajectory and predicts future latent states from which DreamGuard derives immediate-hazard and prefix-risk evidence. It then fuses these multi-horizon signals into intervention decisions before execution. Experiments across four benchmarks and an online guardrail evaluation show that DreamGuard outperforms generic, reactive, and proactive guardrail baselines, achieves the best safety-utility trade-off among evaluated guardrails, and maintains an average end-to-end latency of 25 ms per call.
Wenhao Lin, Chenyu Yu, Xingwei Lin +6
Zhejiang University · Nanjing University of Posts and Telecommunications · Sun Yat-sen University
Large language model (LLM) agents are vulnerable to prompt-injection attacks that propagate through multi-step workflows, tool interactions, and persistent context, making input-output filtering alone insufficient for reliable protection. This paper presents SafeAgent, a runtime security architecture that treats agent safety as a stateful decision problem over evolving interaction trajectories. The proposed design separates execution governance from semantic risk reasoning through two coordinated components: a runtime controller that mediates actions around the agent loop and a context-aware decision core that operates over persistent session state. The core is formalized as a context-aware advanced machine intelligence and instantiated through operators for risk encoding, utility-cost evaluation, consequence modeling, policy arbitration, and state synchronization. Experiments on Agent Security Bench (ASB) and InjecAgent show that SafeAgent consistently improves robustness over baseline and text-level guardrail methods while maintaining competitive benign-task performance. Ablation studies further show that recovery confidence and policy weighting determine distinct safety-utility operating points.
Hailin Liu, Eugene Ilyushin, Jie Ni +1
Lomonosov Moscow State University · Central University
As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat actions in isolation, missing how risks accumulate, while trajectory-level evaluations operate post-hoc, offering no opportunity for timely intervention. To address these limitations, we formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, when to intervene, and what the risk is. We introduce PASTABench, a benchmark of 1,139 multi-turn trajectories spanning 5 risk categories and 13 subcategories. We further propose the Optimal Intervention Window (OIW), anchored by annotated Earliest-Signal and Trigger turns, to quantify intervention timeliness. Evaluation of 16 LLMs reveals that proactive intervention remains largely unsolved, with the best model achieving only 40.74% optimal-timing interventions. Fine-grained diagnosis further uncovers pervasive lexical overfitting: competitive safety scores of smaller models mask keyword hypersensitivity rather than genuine risk comprehension, as their proactive capability largely collapses once hazard vocabulary is neutralized.
Jiapeng Sun, Yujin Zhou, Han Zhu +4
The Hong Kong University of Science and Technology, Hong Kong, China · Peking University, Beijing, China