Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.
Figures & tables
Figure 1: The OCS/ICS-delta pipeline (top), each instrument’s logic (middle), and both applied to one real case (bottom): swapping the rule or negating its obligation flips this model’s verdict, but its compliance-predictive activation moves by no more than 7.9% of baseline magnitude. The cosine/angle notation is a conceptual simplification; Sections 3.1 – 3.2 give the exact definitions used throughout.
Model
OCS-del
OCS-swap
OCS-neg
OCS-agg
Qwen2.5-7B
0.052
0.118
0.135
0.094
Llama-3.1-8B
0.064
0.079
0.060
0.070
Llama-3.3-70B
0.061
0.095
0.084
0.081
Llama-Guard-3-8B
0.014
0.015
0.009
0.013
Mistral-7B-v0.3
0.076
0.103
0.065
0.089
Table 1: OCS by perturbation condition, pooled across all 20 domains (200 cases/domain, seed 42). Per-domain values in Figure 2 .
Figure 2: Domain × model OCS-agg. Nearly every cell is cold (low rule sensitivity); the one visibly warm cell (Mistral-7B-v0.3 × cybersecurity_mitre_attack, 0.265) is the single highest value in the dataset.
Dataset
OCS-agg
Baseline acc.
n
OmniCompliance-100K †
0.081
0.909
3,946
LegalBench
0.255
0.497
887
ContractNLI
0.318
0.786
387
Table 2: Generalization check, Llama-3.3-70B-Instruct only. † OmniCompliance-100K row restated from Table 1 and Table 5 . OCS stays well under the 0.5 independence reference everywhere but is not a fixed, dataset-independent property of the model. n counts sampled cases, except for ContractNLI, where it counts the cases with a parseable baseline verdict.
Model
Part.
−1
−0.5
0
+0.5
+1
Llama-3.1-8B
LOW
92.3 ∗
92.5
92.8
92.5
92.3 ∗
Llama-3.1-8B
HIGH
79.7
82.2
82.2
82.7
82.2
Llama-3.3-70B
LOW
91.7 ∗∗
91.9
92.1
92.3
92.3
Llama-3.3-70B
HIGH
73.3 ∗
76.7
79.2
79.2
78.7
Table 3: Accuracy (%) by steering α , all four (model, partition) cells. Bold marks the unsteered baseline. ∗p<0.05 , ∗∗p<0.01 (McNemar’s exact test vs. the baseline; nominal, uncorrected for the 16 comparisons in this table). Exact p -values for every cell in Appendix H .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
n total
Modal cov.
gdpr
3,759
59.1%
hipaa
6,924
20.5%
eu_ai_act
7,758
60.7%
ccpa
2,489
30.9%
sb35
1,244
31.8%
finance_crypto
1,969
38.9%
Appendix
Table 4: Domain sizes and per-sample mandatory-modal coverage (OCS-neg applicability). The three domains at 0.0% (MITRE ATT&CK technique descriptions, academic-integrity value statements, and online-learning guidance) are not phrased in mandatory-obligation language at all. This is a genuine linguistic property of these domains, not a processing failure; with nothing to negate, they are excluded from OCS-neg entirely. † Domain name as in OmniCompliance-100K.
Model
Baseline
Del
Swap
Neg
Qwen2.5-7B
90.3%
90.7%
83.2%
79.5%
Llama-3.1-8B
91.9%
91.4%
89.4%
89.7%
Llama-3.3-70B
90.9%
87.5%
84.5%
83.6%
Llama-Guard-3-8B
51.1%
50.0%
50.1%
50.9%
Mistral-7B-v0.3
90.3%
91.6%
89.6%
88.7%
Appendix
Table 5: Baseline accuracy and accuracy under each rule perturbation condition, pooled across all 20 domains. Accuracy under rule deletion approximates performance when only case content is available.
Figure 3: Mean OCS by perturbation condition, pooled across all models and domains. OCS-del (rule deleted) is the lowest of the three: verdicts are least likely to change when the rule is removed entirely.
Model
mean ∣ ICS ∣
mean ∣ ICS-delta ∣
Ratio
Llama-3.1-8B
4.56
0.58
12.8%
Llama-3.3-70B
3.31
0.69
21.0%
Mistral-7B-v0.3
1.45
0.33
22.9%
Qwen2.5-7B
16.71
2.79
16.7%
Appendix
Table 6: Mean baseline ICS magnitude, mean ∣ ICS-delta ∣ , and their ratio, pooled across all three perturbation conditions and 20 domains.
Figure 4: ICS-delta as a percentage of baseline ICS magnitude, per model per condition. Every bar is a small fraction of the baseline signal; OCS-neg (rightmost, green) is consistently the smallest.
Model
ρ
p
Qwen2.5-7B-Instruct
0.114
0.631
Llama-3.1-8B-Instruct
0.084
0.724
Mistral-7B-Instruct-v0.3
−0.032
0.895
Llama-3.3-70B-Instruct
0.338
0.145
Appendix
Table 7: OCS-agg vs. ICS-delta-agg per-domain rank correlation, one row per Tier-1 model ( n=20 domains each). None reaches conventional significance.
Model
Null (indep.)
Obs. OCS
Perm. p
Llama-3.1-8B
0.500
0.070
<0.0001
Llama-3.3-70B
0.486
0.081
<0.0001
Mistral-7B-v0.3
0.502
0.089
<0.0001
Qwen2.5-7B
0.483
0.094
<0.0001
Llama-Guard-3-8B
0.019
0.013
<0.0001
Appendix
Table 8: Observed OCS-agg vs. an independence null built from each model’s own empirical baseline/altered verdict marginals ( p(1−q)+q(1−p) ), validated by a 10,000-sample within-domain permutation test ( p : one-sided, observed ≤ null). Every model’s OCS is significantly below its own bias-corrected null, not only below the naive 0.5 reference.
Model
Cond.
n
Obs.
Null
p↓
p↑
Llama-3.1-8B
swap
132
0.780
0.669
1.000
<0.0001
Llama-3.1-8B
neg
42
0.333
0.437
0.081
0.987
Llama-3.3-70B
swap
187
0.733
0.723
0.881
0.329
Llama-3.3-70B
neg
53
0.377
0.479
0.036
0.996
Mistral-7B-v0.3
swap
120
0.750
0.665
0.999
0.006
Mistral-7B-v0.3
neg
36
0.389
0.486
0.166
0.961
Appendix
Table 9: Full rule-necessity control statistics: OCS-swap and OCS-neg per model, restricted to that model’s own rule-necessary subset (baseline correct, del incorrect). Llama-3.3-70B-Instruct’s nominal p↓=0.036 for OCS-neg does not survive any correction for the ten tests conducted here (Bonferroni threshold 0.005 ); we do not treat it as a significant reduction.
Figure 5: OCS-agg vs. baseline accuracy, all 100 (model, domain) cells. The correlation is positive but narrowly misses p<0.05 .
Feature
Spearman ρ
pBH
Sig. (BH 0.05)
ECD
0.041
0.86
no
CRC
0.241
0.38
no
MMR
0.405
0.14
no
MDD
0.570
0.044
yes
FKGL
0.448
0.12
no
Appendix
Table 10: Linguistic feature correlations with per-domain mean OCS-agg ( n=19 –20 domains). Only MDD survives Benjamini-Hochberg correction.
Figure 6: Standardized OLS coefficients (all five features jointly, → OCS-agg). Red bars mark p<0.05 individually within the joint model; none reach this threshold, and MDD is closest.
Model
Δlow
Δhigh
p (LOW)
Llama-3.1-8B
−0.006
0.000
0.146
Llama-3.3-70B
+0.007
−0.033
0.073
Qwen2.5-7B
−0.028
−0.006
<0.0001
Mistral-7B-v0.3
−0.010
0.000
0.077
Llama-Guard-3-8B †
−0.001
0.000
0.500
Appendix
Table 11: Rule-grounded prompting: accuracy change vs. standard prompting on the LOW/HIGH-OCS partitions (Llama-3.3-70B-Instruct’s per-sample OCS), parseable-verdict cases only; p is McNemar’s exact test on LOW. Only Qwen2.5-7B reaches significance, and harmfully. † Robustness check, not a grounded-reasoning test (Section 4.3 ).
Model
Part.
α=−1
−0.5
+0.5
+1
Llama-3.1-8B
LOW
0.0489
0.1628
0.1496
0.0273
Llama-3.1-8B
HIGH
0.2668
1.0000
1.0000
1.0000
Llama-3.3-70B
LOW
0.0044
0.1094
0.1153
0.1742
Llama-3.3-70B
HIGH
0.0290
0.3323
1.0000
1.0000
Appendix
Table 12: Exact McNemar’s-test p -values, each (model, partition, α ) cell vs. that row’s α=0 baseline, paired per case.
Figure 7: Accuracy vs. steering α , LOW- and HIGH-OCS partitions, both models. Negative α degrades accuracy for both models on both partitions; positive α never significantly helps.
Current approaches to AI compliance treat conformity as a binary, audit-time verdict rather than a continuous, measurable property of production systems. We argue that this compliance fiction is structurally ill-suited to the requirements of the EU AI Act, which demands ongoing human oversight and the detection of emergent behavioural drift in deployed systems. We introduce governance from metrics, a principle whereby regulatory compliance is derived as a continuous signal from runtime observability rather than from static assessments. Building on this principle, we present govllm, an open-source framework implementing a governance-driven routing architecture in which model selection is determined by accumulated compliance scores rather than by latency or cost alone. Central to our approach is a panel of regulatory judges - LLM evaluators specialised per criterion (EU AI Act, GDPR, ANSSI, accessibility) - whose inter-judge disagreement we reframe not as noise but as a regulatory uncertainty signal warranting human arbitration. We validate this approach through a ground truth corpus of 49 annotated prompt/response pairs across five regulatory criteria, evaluated by four small language models (SLMs, 1.7B-7B parameters) running fully on-premise. Agreement rates range from 51.5% (mistral:7b) to 69.1% (phi4-mini), with no single model dominating across all criteria - empirically motivating the Profile-as-jury design. We further document three structural failure modes in small regulatory judges and a judge-specific position bias that degrades agreement by up to 25 percentage points across three question-order conditions (original, reversed, permuted). govllm is released as open-source software to support reproducible AI governance research.
Large Language Models (LLMs) are increasingly adopted for compliance and legal reasoning tasks, yet their outputs often lack explicit grounding in legal logic and evidence. We present Code-as-Auditor, an LLM-based framework that extends the model's reasoning capability toward structured and evidence-grounded compliance assessment. The framework translates regulatory information into (1) formalized checklists and executable decision trees, encoding regulations and conditions as interpretable code structures. During inference, each checklist item is (2) dynamically expanded into factual and counterfactual questions, guiding the model to reason over case-specific evidence and potential violations. This process establishes a reasoning pipeline that proceeds from evidence identification, through rule application, to final decision-making, while a self-verification loop improves the logical consistency of the generated code and the traceability of outcomes. Experiments on privacy and data protection scenarios demonstrate that Code-as-Auditor delivers more accurate and evidence-backed evaluations, enabling automated compliance regulation checking grounded in explicit regulatory criteria.
As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulatory penalties and direct consumer harm. Existing guard models are built around general harm taxonomies and overlook violations grounded in specific financial regulations. We address this gap with a regulation-driven pipeline that operates directly on regulatory documents, inducing a financial compliance risk taxonomy and synthesizing grounded training data without any predefined violation categories. Instantiating the pipeline on Chinese financial regulations, we release \textbf{FinGuard-Bench}, to our knowledge the first benchmark for financial regulatory compliance detection, with expert-annotated labels at both the query and response levels. We further train \textbf{FinGuard}, a financial compliance detection model built on Qwen3-8B and trained on the regulation-grounded data via supervised fine-tuning and self-play reinforcement learning. On FinGuard-Bench, FinGuard substantially outperforms all baselines, including dedicated guard models and much larger general-purpose LLMs such as Qwen3.5-397B-A17B and GPT-5.1. Furthermore, FinGuard also preserves general safety capabilities and adapts to unseen institution-specific policies using policy documents alone. We will publicly release the code, prompts, and resources used in this work on GitHub.
Huaixia Dou, Jie Zhu, Minghao Wu +5
Qwen DianJin Team, Alibaba Cloud Computing · Tongyi Lab, Alibaba Group · School of Computer Science and Technology, Soochow University