Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency
Figures & tables
Figure 1: AdaGuard framework takes in a user input along with a list of policies, and generates a Violation or Compliance verdict for each policy. AdaGuard supports three reasoning modes: (1) \rmode reasoning_off: no reasoning for any policy, (2) \rmode reasoning_on: reasoning enforced for all policies, and (3) \rmode autoreasoning: reasoning conditioned on a self-predicted difficulty label.
Model
Aegis 2.0
Harm- Bench
Wild- Guard
Sorry- Bench
XS- Test
Dynamic- Policy (Ours)
Avg Latency (s)
Output Tokens
Tokens / sec
Static Avg
All-Tasks Avg
Open-source Guardrails
WildGuard-7b
83.0
86.0
74.2
58.2
93.2
77.5
1.44
20.7
14.4
78.9
78.7
LlamaGuard-3-8b
71.8
84.2
69.9
59.1
88.8
73.7
0.83
3.4
4.1
74.8
74.6
ShieldGemma-9b
73.7
44.1
41.6
39.0
60.2
71.2
1.32
1.0
0.8
51.7
55.0
DynaGuard-8b (no CoT)
78.9
87.1
79.3
86.5
88.2
79.6
2.34
57.7
24.7
84.0
83.3
DynaGuard-8b (with CoT)
80.5
87.0
80.8
86.2
89.6
79.4
6.73
150.5
22.4
84.8
83.9
Table 1: Per-example F1 (%) on static safety benchmarks and our held-out dynamic-policy test set, with per-example efficiency (average latency, output tokens, decoding speed) and Static / All-Tasks F1 averages. All evaluations are conducted on 4 A100 GPUs. We report example-level scores since baseline models typically output aggregate compliance verdicts across all policies for a given input. For frontier models, we use their default reasoning modes and prompt them to follow our \rmode reasoning_off format. For our AdaGuard model, we calibrate across the three reasoning modes. Bold = best per column, underline = second best (Output tokens is descriptive and not ranked).
Figure 2: \rmode autoreasoning behavior in GRPO training on the dynamic-policy validation set, for (left) Gemma4-E4B and (right) Nemotron-Nano-4B. Each panel plots the fraction of policies the model chooses to reason on, and the percentage of the reasoned blocks in \rmode autoreasoning where the model fixes an incorrect counterfactual verdict ( Fixed ) versus turns a correct one wrong ( Broke ). Fixed exceeds Broke throughout in both models, so reasoning is net-positive at the verdict level.
Gemma4-E4B
Nemotron-Nano-4B
Stage
F1 Score
Autoreasoning Rate
F1 Score
Autoreasoning Rate
Base
74.76
21.9
26.75
56.23
SFT
76.51
0.42
46.31
0
GRPO
80.1
31.81
64.52
36.32
Table 2: Training-stage ablation in \rmode autoreasoning, for Gemma4-E4B and Nemotron-Nano-4B on the dynamic-policy test set. Per-policy F1 is reported per stage. Autoreasoning Rate is the fraction of policies the model reasons for.
Figure 3: Reward-feature ablation on Gemma4-E4B. Left: \rmode autoreasoning reasoning rate during GRPO (validation); the CF_Corr-off variants (Quality Only, Vanilla) collapse to ∼100% while CF_Corr-enabled ones stay selective. Right: held-out per-policy F1 vs. reasoning rate at epoch 1; CF_Corr alone prunes hardest but scores lowest, and adding the quality judge restores always-reasoning F1 at a rate near a third.
Figure 4: Per-policy F1, easy vs. hard teacher splits , by stage (Base / SFT / GRPO) and mode, for (left) Gemma4-E4B and (right) Nemotron-Nano-4B. Base/SFT use forced \rmode reasoning_off / \rmode reasoning_on; GRPO uses \rmode autoreasoning. Easy F1 is near-flat, so the hard split separates stages and modes.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Early-stopping patience
8 (eval loss)
Per-device train batch size
2
Gradient accumulation steps
2
Weight decay
0.1
Warmup ratio
0.1
Gradient checkpointing
enabled
Appendix
Table 3: SFT hyperparameters for the best-performing configuration identified in the SFT sweep (Section 2.3.1 ). Values are read directly from the internal training-configuration tracking sheet for that configuration.
Hyperparameter
Value
Training & rollout
Algorithm
\rmode gdpo
Loss mode
\rmode vanilla
Learning rate
5×10−6
LR scheduler
constant
Warmup steps
16
Appendix
Table 4: GRPO hyperparameters for the \rmode autoreasoning configuration, which enables the counterfactual reasoning-value reward (CF_Corr). Reward constants (bottom block) are defined in Appendix A.10 .
Criterion
Accuracy
F1
Gwet’s AC1
Logic
0.878
0.933
0.855
Informativeness
0.854
0.918
0.818
Policy Structure
0.915
0.953
0.898
Policy Boundaries
0.854
0.919
0.822
Final Response Correctness
0.927
0.961
0.915
Mean across criteria
0.885
0.937
0.862
Appendix
Table 5: Evaluator-vs-human-annotator agreement on 82 Gemma-4-31B -generated reasoning traces, evaluated by GPT-OSS-120B (automated evaluator) against a human annotator’s independent scoring, one row per rubric criterion (Appendix A.5 , Stage 2). Values computed directly from the recorded per-criterion pass/fail calls; no estimation.
Split
# inputs
Mean K
Mean violation ratio
Mean difficulty ratio
Unique categories
SFT train
13,008
5.05
0.164
0.155
112
SFT val
6,038
5.44
0.163
0.137
111
GRPO train
3,500
6.29
0.180
0.327
100
GRPO val
1,952
6.31
0.165
0.339
87
Holdout (dynamic-policy test)
12,291
6.61
0.187
0.307
15
Total (all splits, deduplicated)
75,254
—
—
—
127
Appendix
Table 6: Released training-corpus split statistics. Violation ratio is ∣{k:vk=\textscViolation}∣/K averaged per input (the L1 quantity of Appendix A.6 ). While the distribution of datapoints with no input-policy pair violations to datapoints with at least one violation is split uniformly across datasets, this ratio measures the ratio of per-input-policy-pair violations. Per-policy difficulty dk is not present in this export, so target-vs-achieved L1–L4 values await the final curation config (flagged in Appendix A.6 ). “Unique categories” counts distinct input-category values within that split; the five training/validation splits share 0 categories with the holdout split (Table 7 ).
Holdout-only input category
Child Sexual Abuse & Exploitation
Criminal Planning & Non-Violent Crime
Disinformation & Misinformation
Drugs & Controlled Substances
Fraud, Deception & Economic Harm
Harassment, Bullying & Toxic Language
Appendix
Table 7: The 15 final input-category values encompassing the holdout set (no exact overlap with SFT/GRPO train or validation split categories), supporting the zero-shot dynamic-policy generalization claim.
Figure 5: GRPO dataset curation. The refiner successfully flattens extreme left-skewed (zero-inflated) distributions in L1 and L2, pulling the dataset toward the μ=0.30 target. L3 constraints ensure high-difficulty rates are balanced across both compliance and violation labels.
Figure 6: Diagnostics for GRPO pruning. The top plot shows how the number of matched policies per input is reduced after pruning in order to achieve desired difficulty and violation ratios, while the L4 bar chart shows individual semantic categories adhering to the global difficulty target.
Figure 7: SFT dataset curation. The refiner mitigates severe skewness in the raw data, aligning violations ( μ=0.30 ) and difficulty ( μ≈0.17 ) properly within tolerance limits.
Figure 8: Diagnostics for SFT pruning. The top plot shows how the number of matched policies per input is reduced after pruning in order to achieve desired difficulty and violation ratios, while the L4 bar chart shows individual semantic categories adhering to the global difficulty target.
Figure 9: An additional round of pruning on SFT Train/Dev and GRPO Train/Dev, defined by a shift of difficult datapoints with a high matched number of policies from SFT into GRPO training sets. This was done with the purpose of (1) increasing the matched number of policies in GRPO sets, and (2) enforcing higher difficulty in GRPO sets.
Figure 10: Per-category F1 score breakdown across eight safety taxonomies comparing BASE, SFT, and RL AdaGuard models.
#
Scenario (CF-CORR branch)
Base rc
Good quality
Bad quality
1
Flip (own correct, off wrong)
+1.5
+2.25
+0.75
2
Broke (own wrong, off correct)
−1.25
−1.25 (clamped)
−1.875
3
Wrong-both (own wrong, off wrong)
−0.75
−0.75 (clamped)
−1.125
4
Unnecessary, teacher High
+1.0
+1.0 (capped; +1.5 pre-fix)
+0.5
5
Unnecessary, teacher Low
+0.8
+0.8 (capped; +1.2 pre-fix)
+0.4
6
Predicted- Low , correct (never judged)
+1.0
+1.0
+1.0
Appendix
Table 8: Every CF-Corr scenario (Eq. 16 ) under the sign-aware, unnecessary -aware quality multiplier of Eq. 6 , at good quality ( q=5⇒m=1.5 ) and bad quality ( q=0⇒m=0.5 ). Rows 1–3 use the ordinary branches ( m on correct, max(2−m,1) on wrong); rows 4–5 are the unnecessary cell and use the clamp min(m,1.0) ; rows 6–7 emit no trace, are never judged, and receive no multiplier. The invariant that closes the amplification pathway is visible in the good-quality column: rows 4–5 ( +1.0 , +0.8 ) never exceed the correctly-skipped baseline of row 6 ( +1.0 ), whereas the pre-fix unclamped m(q) would have lifted them to +1.5 and +1.2 .