The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.
Figures & tables
Figure 1: The PROACT-Agent ecosystem. (a) PROACT-Agent formulates risk as a non-Markovian prefix-evaluation task, allowing intervention on observed context before the next LLM inference. (b) The framework combines trajectory context, prefix-level evaluation, and monotonic causal rectification. (c) PROACT-Bench contains 155,780 bilingual states with available or inferred auxiliary rationales and model-adjudicated safety labels.
Figure 2: Overview of the PROACT-Agent framework, which transforms static logs into a causally-rigorous streaming dataset across three stages: (1) Progressive Trajectory Unrolling (PTU) unrolls trajectories into non-Markovian evaluation states; (2) Reasoning-Augmented Causal Rectification (RACR) infers missing auxiliary rationales and rectifies labels via Monotonic Causal Consistency (MCC); and (3) Culturally-Aware Data Localization (CADL) maps states across English and Chinese contexts with Tripartite Rigid Verification (TRV) checks.
Figure 3: Overview of the PROACT-Bench dataset. (a) Scale and language: 155,780 prefix-level states across English and localized Chinese environments. (b) Quality verification: Data Quality Score (DQS) summarizes model-committee agreement and contested labels within the adjudication pipeline. (c) Taxonomic coverage: the benchmark maps six safety risk categories across a range of operational environments.
Method
Overall
English Subset
Chinese Subset
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Closed-source LLMs
Qwen3-Max [ 1 ]
87.52
80.25
98.25
88.34
87.04
78.66
98.62
87.52
88.00
81.80
97.90
89.13
Gemini 3 Flash [ 8 ]
88.73
84.30
94.22
88.98
88.08
83.15
92.98
87.79
89.38
85.35
95.34
90.07
Open-source LLMs
DeepSeek-V3.2 [ 19 ]
86.24
79.77
95.74
87.03
81.55
73.12
94.84
82.57
90.96
86.93
96.58
91.50
Table 1: Broad Mixed-Source Comparison. Accuracy (Acc.), unsafe Precision (Pre.), Recall (Rec.), and F1 (%) on the Overall, English, and Chinese subsets of the mixed-source comparison protocol. Best results are bold ; second-best results are underlined .
Strict Root-Grouped
Complete Source-Held-Out
Method
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Qwen3Guard-Gen-8B
94.67
88.59
93.96
91.20
75.09
76.67
66.77
71.38
TS-Guard
90.36
79.40
90.75
84.70
75.06
74.81
69.94
72.29
PROACT-Agent
97.85
96.44
96.23
96.34
92.37
95.52
87.72
91.46
Table 2: Strict Generalization. Acc., unsafe Pre., Rec., and F1 (%) on 31,156 root-grouped and 3,396 complete source-held-out evaluation states. Invalid outputs count as errors.
Strict Root-Grouped
Complete Source-Held-Out
Method
First-Unsafe
Exact
First-Unsafe
Exact
Qwen3Guard-Gen-8B
94.39
93.18
62.93
57.99
TS-Guard
91.05
90.14
64.67
60.42
PROACT-Agent
97.38
97.12
91.84
90.63
Table 3: Temporal Boundary Detection (%). First-unsafe recall requires detecting the gold first-unsafe prefix as unsafe; exact-boundary detection additionally requires every earlier safe prefix to remain unblocked. The denominator is root-language trajectories with an unsafe boundary, with EN/ZH evaluated separately. Gold boundaries are model-adjudicated; invalid predictions count as errors.
Defense
Successful Attacks ↓
Banking ASR ↓
Slack ASR ↓
Travel ASR ↓
Workspace ASR ↓
Overall ASR ↓
No guard
2,173 / 10,439
36.87%
56.62%
40.06%
5.16%
20.82%
+ PROACT-Agent
42 / 10,439
0.00%
0.00%
2.73%
0.00%
0.40%
Table 4: AgentDojo Closed-Loop Results on Non-DoS Attacks. Suite-wise and overall targeted attack success rates (ASR) across four suites. Each paired logical case is identified by suite, attack, user task, and injection task; the overall evaluation contains 10,439 cases.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Fixed ID 20%
Complete OOD
Training
Selected Step
Macro-F1 ↑
FBR ↓
FAR ↓
Macro-F1 ↑
FBR ↓
FAR ↓
60%
8,500
98.08
0.96
3.10
90.55
3.69
15.76
80%
8,000
98.24
0.91
2.79
91.78
4.13
12.72
Δ (80 − 60)
—
+0.16
−0.05
−0.31
+1.23
+0.44
−3.04
Appendix
Table 6: Matched Two-Source Training-Size Sensitivity. Selected checkpoints from AgentAlign+ToolSafety training runs share fixed-ID (30,470 states) and complete-OOD (3,396 states) evaluations. Metrics are percentages; Δ denotes 80% minus 60% in percentage points. Invalid predictions are zero in all settings.
Backbone
Acc.
Pre.
Rec.
F1
Invalid
Qwen3-1.7B
96.14
95.30
96.73
96.01
0
Qwen3-4B
96.93
96.51
97.11
96.81
0
Qwen3-8B
96.91
96.77
96.80
96.78
4
Appendix
Table 7: Qwen3 Backbone-Scale Robustness. Accuracy, unsafe Precision, Recall, and F1 (%) on the same mixed-source evaluation set for selected Qwen3 checkpoints. Invalid predictions count as errors.
Backbone
Acc.
Pre.
Rec.
F1
Invalid
Qwen2.5-7B-Instruct
97.11
96.30
97.74
97.01
0
Qwen3-8B
96.91
96.77
96.80
96.78
4
Qwen3.5-9B
97.16
96.85
97.25
97.05
0
Appendix
Table 8: Cross-Generation Backbone Robustness. Accuracy, unsafe Precision, Recall, and F1 (%) on the same mixed-source evaluation set. Model generation and parameter count both vary, so this is not a controlled scaling experiment. Qwen3-8B intentionally bridges this comparison and Table 7 .
PROACT backbone
Standard P50 / P95
Memory (GiB)
Slope / 1K (ms)
4K+ P50 / P95
6K+ P50 / P95
Qwen3-1.7B
403 / 889
4.69
20.8
575 / 1,138
657 / 1,151
Qwen3-4B
543 / 1,137
8.78
62.1
894 / 1,714
1,118 / 1,723
Qwen2.5-7B-Instruct
474 / 976
15.63
86.6
899.8 / 1,629
1,159 / 1,633
Appendix
Table 9: Guard-Only Runtime Efficiency and Long-Context Scaling. Measurements use one A800-SXM4-80GB at batch size 1. Standard latency summarizes the overall guard-only workload; 4K+ and 6K+ summarize long-context subsets with input lengths exceeding the corresponding token thresholds. Slope is the empirical latency-growth trend per additional 1K input tokens. Latency is in milliseconds; memory is P95 peak allocated memory in GiB.
Method
Overall
English Subset
Chinese Subset
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
Acc.
Pre.
Rec.
F1
W/o CADL
91.19
85.89
97.79
91.45
90.21
83.23
98.62
90.27
92.18
88.54
97.02
92.59
W/o RACR
92.09
91.65
91.98
91.82
91.57
87.83
94.84
91.20
92.62
95.72
89.35
92.42
W/o PTU
93.48
93.67
92.74
93.20
92.52
90.12
94.09
92.06
94.44
97.30
91.50
94.31
PROACT-Agent (Full)
94.53
93.16
95.68
94.40
93.57
89.56
97.39
93.31
95.50
96.87
94.10
95.46
Appendix
Table 10: Mixed-Source Ablation. Impact of independently removing the three core operators (CADL, RACR, and PTU) from the PROACT-Agent framework under the mixed-source comparison protocol. All variants are trained using the Qwen2.5-7B [ 35 ] backbone. The full pipeline achieves the best Accuracy and F1-score across all evaluation cohorts, while maintaining a strong Precision–Recall balance. Unsafe is treated as the positive class. Best results are shown in bold and second-best results are underlined .
Figure 4: The system prompt deployed to the cognitive parser ( Mparse ), which is used in the PTU module.
Figure 5: The system prompt for reasoning recovery used in the RACR module.
Figure 6: The system prompt for the diverse evaluator committee ( M∈M ) used in the RACR module - Part 1: Persona definition and baseline safety rules.
Figure 7: The system prompt for the diverse evaluator committee ( M∈M ) used in the RACR module - Part 2: High-risk categorization and output formatting constraints.
Figure 8: The system prompt for the information-augmented arbitrator ( J ) used in the RACR module.
Figure 9: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 1: Persona definition, core principles, and detailed mapping rules.
Figure 10: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 2: An abridged few-shot English baseline example reproduced from the localization template.
Figure 11: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 3: Explicit rationale formulation demonstrating how entities and risks are mapped onto the Chinese cultural manifold.
Figure 12: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 4: An abridged localized JSON few-shot example reproduced from the localization template.
Figure 13: The Chinese-authored system instruction for the generative localization model used in the CADL module - Part 5: Dynamic data injection and the constrained two-step output workflow. Template placeholders are rendered before model invocation.