Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard
Figures & tables
Figure 1: Policy-adaptive assessment of an agent trajectory. The request, tool call, and successful delivery are identical in both cases; only the governing policy changes. The guard jointly considers the supplied policy and the recorded interaction, generates an overall analysis, and outputs the violated rule identifiers or NR when no rule is violated. The analysis is generated by the guard and does not represent the assessed agent’s internal reasoning.
Figure 2: Overview of AdaGuard. AdaptiveSafety supports supervised initialization through variation in policies and agent behavior. SafePO evaluates complete responses using structured verdict rewards and distributes their group-relative learning signals through value-guided token weights. Separate normalization preserves fixed analysis and verdict weight budgets. The value model and frozen reference are used only during training.
AdaptiveSafety
DynaBench
Model
Acc.
Prec.
Rec.
F1
Acc.
Prec.
Rec.
F1
Qwen3Guard-Gen-4B
59.00
74.46
27.40
40.06
51.57
83.33
1.87
3.66
Qwen3Guard-Gen-8B
58.90
75.72
26.20
38.93
51.93
100.00
2.25
4.40
YuFeng-XGuard-Reason-8B
57.30
66.82
29.00
40.45
50.09
48.36
22.10
30.33
DynaGuard-4B
67.30
67.76
66.00
66.87
74.22
77.25
67.42
72.00
DynaGuard-8B
64.10
61.97
73.00
67.03
80.11
88.04
68.91
77.31
Table 1: Binary assessment of local models (%). Violations are positive and invalid outputs count as errors. Bold marks column maxima. Interface and truncation details are in Appendix A .
AdaptiveSafety
DynaBench
Model
Exact
Rule-F1
Exact
Rule-F1
Laya
31.80
6.27
49.91
3.38
Laya-Typed-Decisions
8.50
8.98
49.36
2.05
AdaGuard-0.6B
68.10
59.28
44.94
50.50
AdaGuard-4B
76.60
71.50
63.72
54.99
AdaGuard-8B
77.10
74.05
70.72
63.42
Table 2: Rule identification by local models with policy-local outputs (%). Exact is complete-set accuracy; Rule-F1 is micro-F1. Binary-only interfaces are excluded. Full rule precision and recall appear in Table A2 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
AdaptiveSafety
DynaBench
Model
Acc.
Prec.
Rec.
F1
Acc.
Prec.
Rec.
F1
GPT-5.2
76.10
69.28
93.80
79.69
83.06
88.89
74.91
81.30
Gemini-3.1-Flash-Lite
72.10
75.52
65.40
70.10
80.85
80.07
81.27
80.67
Jev
82.40
79.03
88.20
83.36
83.61
93.20
71.91
81.18
Qwen3Guard-Gen-4B
59.00
74.46
27.40
40.06
51.57
83.33
1.87
3.66
Qwen3Guard-Gen-8B
58.90
75.72
26.20
38.93
51.93
100.00
2.25
4.40
Appendix
Table A1: Binary safety assessment (%). Unsafe is the positive class; invalid outputs count as errors. Bold marks the highest value in each column among the reported runs. Laya results use native SDK truncation.
AdaptiveSafety
DynaBench
Model
Exact
Micro-P
Micro-R
Micro-F1
Exact
Micro-P
Micro-R
Micro-F1
GPT-5.2
52.70
44.17
78.18
56.45
82.87
88.50
74.91
81.14
Gemini-3.1-Flash-Lite
61.00
64.80
42.31
51.20
80.85
80.07
81.27
80.67
Jev
57.60
39.70
74.82
51.88
80.85
86.82
71.54
78.44
Laya
31.80
3.34
50.07
6.27
49.91
6.82
2.25
3.38
Laya-Typed-Decisions
8.50
4.86
59.59
8.98
49.36
12.00
1.12
2.05
Appendix
Table A2: Policy-local rule identification (%). Exact requires the complete valid violation set; micro scores pool rule decisions within each sample before aggregation. Failed predictions contribute an empty set to micro counts and never count as exact. Binary-only adapters are omitted. Bold marks the column maximum.
AdaptiveSafety
DynaBench
Model
Err.
Trunc.
M-P
M-R
M-F1
Err.
Trunc.
M-P
M-R
M-F1
GPT-5.2
2
–
79.84
76.10
75.33
0
–
83.91
82.92
82.91
Gemini-3.1-Flash-Lite
103
–
72.50
72.10
71.97
0
–
80.85
80.85
80.85
Jev
0
–
82.84
82.40
82.34
0
–
85.47
83.42
83.33
Qwen3Guard-Gen-4B
0
0.00
64.99
59.00
54.45
0
0.00
67.27
50.76
35.66
Qwen3Guard-Gen-8B
0
0.00
65.55
58.90
53.98
0
0.00
75.70
51.12
36.15
Appendix
Table A3: Coverage and macro metrics (%). Err. is the number of invalid or failed responses. Trunc. is the fraction with detected input truncation; – means unknown. M-P, M-R and M-F1 denote binary macro averages.
AdaptiveSafety
DynaBench
Size
Stage
Acc.
F1
Exact
Acc.
F1
Exact
0.6B
SFT
81.70
80.09
66.70
51.01
62.32
43.83
0.6B
SafePO
82.60
80.71
68.10
51.38
60.71
44.94
4B
SFT
89.90
89.22
77.20
72.56
73.15
63.35
4B
SafePO
89.30
88.53
76.60
71.82
71.82
63.72
8B
SFT
88.40
87.55
76.20
74.77
74.49
67.40
Appendix
Table A4: Historical SFT and current SafePO checkpoints (%). These are descriptive comparisons, not controlled component ablations. The input files match, but the historical preprocessing budget depended on the target length.
Query (117)
Trajectory (883)
Model
Acc.
F1
Exact
Acc.
F1
Exact
GPT-5.2
81.20
75.56
71.79
75.42
80.04
50.17
Gemini-3.1-Flash-Lite
70.09
60.67
61.54
72.37
71.09
60.93
Jev
82.91
71.43
70.94
82.33
84.21
55.83
Laya
65.81
39.39
56.41
52.89
52.51
28.54
Laya-Typed-Decisions
41.03
48.89
15.38
53.79
68.08
7.59
Appendix
Table A5: AdaptiveSafety by assessment scope (%). Query-only inputs assess the request; trajectory inputs assess recorded agent behavior.
Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross-modal composition, while moderation policies may vary across products, regions, and deployment stages. Most existing guardrails either rely on fixed taxonomies or target only a narrow set of interaction settings, which limits their adaptability when safety rules change at deployment time. We present \textbf{SingGuard}, a policy-adaptive multimodal guardrail model family for safety assessment in multimodal conversations. SingGuard treats the active policy as a runtime input: given natural-language rules, it checks the target content against the active policy rule by rule and predicts both the safety label and the triggered rule. To balance efficiency and interpretability, SingGuard supports fast, hybrid, and slow inference regimes along a fast-to-slow reasoning spectrum, ranging from direct safety judgments to policy-grounded deliberation. We further optimize this behavior with fast--slow decoupled reinforcement learning. We also introduce \textbf{SingGuard-Bench}, a multimodal guardrail benchmark with 56{,}340 examples spanning 80+ fine-grained risk types across multimodal QA, adversarial attack, and dynamic-rule evaluation settings, including cross-modal joint-risk cases where each modality is harmless in isolation but their composition implies unsafe intent. Across six benchmark families (35 datasets), SingGuard achieves state-of-the-art average F1 in every family. Dynamic-rule evaluation further shows improved policy-following accuracy from 0.6465 to 0.7415 under runtime policy shifts. Our code is available at https://github.com/inclusionAI/Sing-Guard.
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.
Mingyang Song, Luxin Xu, Haoyu Sun +3
1Fudan University · 2Tongji University · 3Virtue AI +3
Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, safety policies are increasingly specified at inference time by users, organizations, or regulatory contexts. This makes safety enforcement fundamentally dynamic: the guardrail should adapt to changing safety policies without retraining. Yet this requirement creates a fundamental tension: faithfully judging complex policy contexts demands reasoning capability, while practical deployment requires low-latency responses. We introduce Latent Policy Guardrail (LPG), a guardrail framework that learnssemantic latent deliberation over dynamic policies. LPG compresses the internal deliberation needed for intent interpretation and policy grounding into continuous states supervised by decision-relevant semantics. At inference time, it generates only a compact verdict anchored to the violated policy clauses, preserving auditability while avoiding the latency of explicit reasoning. Across policy guardrail benchmarks, LPG-4B reaches 84.5% average safety accuracy and 77.9% F1 by compressing deliberation into just 10 latent tokens, outperforming the strongest dynamic baseline while running roughly 11 times faster than Qwen3-4B-Thinking under the single-sample evaluation setup. Code and data are available at https://github.com/SaFo-Lab/Latent_Policy_Guard.