Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency
Figures & tables
Figure 1: AdaGuard framework takes in a user input along with a list of policies, and generates a Violation or Compliance verdict for each policy. AdaGuard supports three reasoning modes: (1) \rmode reasoning_off: no reasoning for any policy, (2) \rmode reasoning_on: reasoning enforced for all policies, and (3) \rmode autoreasoning: reasoning conditioned on a self-predicted difficulty label.
Model
Aegis 2.0
Harm- Bench
Wild- Guard
Sorry- Bench
XS- Test
Dynamic- Policy (Ours)
Avg Latency (s)
Output Tokens
Tokens / sec
Static Avg
All-Tasks Avg
Open-source Guardrails
WildGuard-7b
83.0
86.0
74.2
58.2
93.2
77.5
1.44
20.7
14.4
78.9
78.7
LlamaGuard-3-8b
71.8
84.2
69.9
59.1
88.8
73.7
0.83
3.4
4.1
74.8
74.6
ShieldGemma-9b
73.7
44.1
41.6
39.0
60.2
71.2
1.32
1.0
0.8
51.7
55.0
DynaGuard-8b (no CoT)
78.9
87.1
79.3
86.5
88.2
79.6
2.34
57.7
24.7
84.0
83.3
DynaGuard-8b (with CoT)
80.5
87.0
80.8
86.2
89.6
79.4
6.73
150.5
22.4
84.8
83.9
Table 1: Per-example F1 (%) on static safety benchmarks and our held-out dynamic-policy test set, with per-example efficiency (average latency, output tokens, decoding speed) and Static / All-Tasks F1 averages. All evaluations are conducted on 4 A100 GPUs. We report example-level scores since baseline models typically output aggregate compliance verdicts across all policies for a given input. For frontier models, we use their default reasoning modes and prompt them to follow our \rmode reasoning_off format. For our AdaGuard model, we calibrate across the three reasoning modes. Bold = best per column, underline = second best (Output tokens is descriptive and not ranked).
Figure 2: \rmode autoreasoning behavior in GRPO training on the dynamic-policy validation set, for (left) Gemma4-E4B and (right) Nemotron-Nano-4B. Each panel plots the fraction of policies the model chooses to reason on, and the percentage of the reasoned blocks in \rmode autoreasoning where the model fixes an incorrect counterfactual verdict ( Fixed ) versus turns a correct one wrong ( Broke ). Fixed exceeds Broke throughout in both models, so reasoning is net-positive at the verdict level.
Gemma4-E4B
Nemotron-Nano-4B
Stage
F1 Score
Autoreasoning Rate
F1 Score
Autoreasoning Rate
Base
74.76
21.9
26.75
56.23
SFT
76.51
0.42
46.31
0
GRPO
80.1
31.81
64.52
36.32
Table 2: Training-stage ablation in \rmode autoreasoning, for Gemma4-E4B and Nemotron-Nano-4B on the dynamic-policy test set. Per-policy F1 is reported per stage. Autoreasoning Rate is the fraction of policies the model reasons for.
Figure 3: Reward-feature ablation on Gemma4-E4B. Left: \rmode autoreasoning reasoning rate during GRPO (validation); the CF_Corr-off variants (Quality Only, Vanilla) collapse to ∼100% while CF_Corr-enabled ones stay selective. Right: held-out per-policy F1 vs. reasoning rate at epoch 1; CF_Corr alone prunes hardest but scores lowest, and adding the quality judge restores always-reasoning F1 at a rate near a third.
Figure 4: Per-policy F1, easy vs. hard teacher splits , by stage (Base / SFT / GRPO) and mode, for (left) Gemma4-E4B and (right) Nemotron-Nano-4B. Base/SFT use forced \rmode reasoning_off / \rmode reasoning_on; GRPO uses \rmode autoreasoning. Easy F1 is near-flat, so the hard split separates stages and modes.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Early-stopping patience
8 (eval loss)
Per-device train batch size
2
Gradient accumulation steps
2
Weight decay
0.1
Warmup ratio
0.1
Gradient checkpointing
enabled
Appendix
Table 3: SFT hyperparameters for the best-performing configuration identified in the SFT sweep (Section 2.3.1 ). Values are read directly from the internal training-configuration tracking sheet for that configuration.
Hyperparameter
Value
Training & rollout
Algorithm
\rmode gdpo
Loss mode
\rmode vanilla
Learning rate
5×10−6
LR scheduler
constant
Warmup steps
16
Appendix
Table 4: GRPO hyperparameters for the \rmode autoreasoning configuration, which enables the counterfactual reasoning-value reward (CF_Corr). Reward constants (bottom block) are defined in Appendix A.10 .
Criterion
Accuracy
F1
Gwet’s AC1
Logic
0.878
0.933
0.855
Informativeness
0.854
0.918
0.818
Policy Structure
0.915
0.953
0.898
Policy Boundaries
0.854
0.919
0.822
Final Response Correctness
0.927
0.961
0.915
Mean across criteria
0.885
0.937
0.862
Appendix
Table 5: Evaluator-vs-human-annotator agreement on 82 Gemma-4-31B -generated reasoning traces, evaluated by GPT-OSS-120B (automated evaluator) against a human annotator’s independent scoring, one row per rubric criterion (Appendix A.5 , Stage 2). Values computed directly from the recorded per-criterion pass/fail calls; no estimation.
Split
# inputs
Mean K
Mean violation ratio
Mean difficulty ratio
Unique categories
SFT train
13,008
5.05
0.164
0.155
112
SFT val
6,038
5.44
0.163
0.137
111
GRPO train
3,500
6.29
0.180
0.327
100
GRPO val
1,952
6.31
0.165
0.339
87
Holdout (dynamic-policy test)
12,291
6.61
0.187
0.307
15
Total (all splits, deduplicated)
75,254
—
—
—
127
Appendix
Table 6: Released training-corpus split statistics. Violation ratio is ∣{k:vk=\textscViolation}∣/K averaged per input (the L1 quantity of Appendix A.6 ). While the distribution of datapoints with no input-policy pair violations to datapoints with at least one violation is split uniformly across datasets, this ratio measures the ratio of per-input-policy-pair violations. Per-policy difficulty dk is not present in this export, so target-vs-achieved L1–L4 values await the final curation config (flagged in Appendix A.6 ). “Unique categories” counts distinct input-category values within that split; the five training/validation splits share 0 categories with the holdout split (Table 7 ).
Holdout-only input category
Child Sexual Abuse & Exploitation
Criminal Planning & Non-Violent Crime
Disinformation & Misinformation
Drugs & Controlled Substances
Fraud, Deception & Economic Harm
Harassment, Bullying & Toxic Language
Appendix
Table 7: The 15 final input-category values encompassing the holdout set (no exact overlap with SFT/GRPO train or validation split categories), supporting the zero-shot dynamic-policy generalization claim.
Figure 5: GRPO dataset curation. The refiner successfully flattens extreme left-skewed (zero-inflated) distributions in L1 and L2, pulling the dataset toward the μ=0.30 target. L3 constraints ensure high-difficulty rates are balanced across both compliance and violation labels.
Figure 6: Diagnostics for GRPO pruning. The top plot shows how the number of matched policies per input is reduced after pruning in order to achieve desired difficulty and violation ratios, while the L4 bar chart shows individual semantic categories adhering to the global difficulty target.
Figure 7: SFT dataset curation. The refiner mitigates severe skewness in the raw data, aligning violations ( μ=0.30 ) and difficulty ( μ≈0.17 ) properly within tolerance limits.
Figure 8: Diagnostics for SFT pruning. The top plot shows how the number of matched policies per input is reduced after pruning in order to achieve desired difficulty and violation ratios, while the L4 bar chart shows individual semantic categories adhering to the global difficulty target.
Figure 9: An additional round of pruning on SFT Train/Dev and GRPO Train/Dev, defined by a shift of difficult datapoints with a high matched number of policies from SFT into GRPO training sets. This was done with the purpose of (1) increasing the matched number of policies in GRPO sets, and (2) enforcing higher difficulty in GRPO sets.
Figure 10: Per-category F1 score breakdown across eight safety taxonomies comparing BASE, SFT, and RL AdaGuard models.
#
Scenario (CF-CORR branch)
Base rc
Good quality
Bad quality
1
Flip (own correct, off wrong)
+1.5
+2.25
+0.75
2
Broke (own wrong, off correct)
−1.25
−1.25 (clamped)
−1.875
3
Wrong-both (own wrong, off wrong)
−0.75
−0.75 (clamped)
−1.125
4
Unnecessary, teacher High
+1.0
+1.0 (capped; +1.5 pre-fix)
+0.5
5
Unnecessary, teacher Low
+0.8
+0.8 (capped; +1.2 pre-fix)
+0.4
6
Predicted- Low , correct (never judged)
+1.0
+1.0
+1.0
Appendix
Table 8: Every CF-Corr scenario (Eq. 16 ) under the sign-aware, unnecessary -aware quality multiplier of Eq. 6 , at good quality ( q=5⇒m=1.5 ) and bad quality ( q=0⇒m=0.5 ). Rows 1–3 use the ordinary branches ( m on correct, max(2−m,1) on wrong); rows 4–5 are the unnecessary cell and use the clamp min(m,1.0) ; rows 6–7 emit no trace, are never judged, and receive no multiplier. The invariant that closes the amplification pathway is visible in the good-quality column: rows 4–5 ( +1.0 , +0.8 ) never exceed the correctly-skipped baseline of row 6 ( +1.0 ), whereas the pre-fix unclamped m(q) would have lifted them to +1.5 and +1.2 .
Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, safety policies are increasingly specified at inference time by users, organizations, or regulatory contexts. This makes safety enforcement fundamentally dynamic: the guardrail should adapt to changing safety policies without retraining. Yet this requirement creates a fundamental tension: faithfully judging complex policy contexts demands reasoning capability, while practical deployment requires low-latency responses. We introduce Latent Policy Guardrail (LPG), a guardrail framework that learnssemantic latent deliberation over dynamic policies. LPG compresses the internal deliberation needed for intent interpretation and policy grounding into continuous states supervised by decision-relevant semantics. At inference time, it generates only a compact verdict anchored to the violated policy clauses, preserving auditability while avoiding the latency of explicit reasoning. Across policy guardrail benchmarks, LPG-4B reaches 84.5% average safety accuracy and 77.9% F1 by compressing deliberation into just 10 latent tokens, outperforming the strongest dynamic baseline while running roughly 11 times faster than Qwen3-4B-Thinking under the single-sample evaluation setup. Code and data are available at https://github.com/SaFo-Lab/Latent_Policy_Guard.
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard
Yunhao Feng, Yifan Ding, Yuxiang Xie +4
National University of Defense Technology · Fudan University
As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained filtering and support fine-grained, interpretable, and adaptable risk assessment. However, existing solutions often rely on rapid classification schemes or post-hoc rules, resulting in limited transparency, inflexible policies, or prohibitive inference costs. To this end, we present YuFeng-XGuard, a reasoning-centric guardrail model family designed to perform multi-dimensional risk perception for LLM interactions. Instead of producing opaque binary judgments, YuFeng-XGuard generates structured risk predictions, including explicit risk categories and configurable confidence scores, accompanied by natural language explanations that expose the underlying reasoning process. This formulation enables safety decisions that are both actionable and interpretable. To balance decision latency and explanatory depth, we adopt a tiered inference paradigm that performs an initial risk decision based on the first decoded token, while preserving ondemand explanatory reasoning when required. In addition, we introduce a dynamic policy mechanism that decouples risk perception from policy enforcement, allowing safety policies to be adjusted without model retraining. Extensive experiments on a diverse set of public safety benchmarks demonstrate that YuFeng-XGuard achieves stateof-the-art performance while maintaining strong efficiency-efficacy trade-offs. We release YuFeng-XGuard as an open model family, including both a full-capacity variant and a lightweight version, to support a wide range of deployment scenarios.