Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.
Figures & tables
Figure 1: The OCS/ICS-delta pipeline (top), each instrument’s logic (middle), and both applied to one real case (bottom): swapping the rule or negating its obligation flips this model’s verdict, but its compliance-predictive activation moves by no more than 7.9% of baseline magnitude. The cosine/angle notation is a conceptual simplification; Sections 3.1 – 3.2 give the exact definitions used throughout.
Model
OCS-del
OCS-swap
OCS-neg
OCS-agg
Qwen2.5-7B
0.052
0.118
0.135
0.094
Llama-3.1-8B
0.064
0.079
0.060
0.070
Llama-3.3-70B
0.061
0.095
0.084
0.081
Llama-Guard-3-8B
0.014
0.015
0.009
0.013
Mistral-7B-v0.3
0.076
0.103
0.065
0.089
Table 1: OCS by perturbation condition, pooled across all 20 domains (200 cases/domain, seed 42). Per-domain values in Figure 2 .
Figure 2: Domain × model OCS-agg. Nearly every cell is cold (low rule sensitivity); the one visibly warm cell (Mistral-7B-v0.3 × cybersecurity_mitre_attack, 0.265) is the single highest value in the dataset.
Dataset
OCS-agg
Baseline acc.
n
OmniCompliance-100K †
0.081
0.909
3,946
LegalBench
0.255
0.497
887
ContractNLI
0.318
0.786
387
Table 2: Generalization check, Llama-3.3-70B-Instruct only. † OmniCompliance-100K row restated from Table 1 and Table 5 . OCS stays well under the 0.5 independence reference everywhere but is not a fixed, dataset-independent property of the model. n counts sampled cases, except for ContractNLI, where it counts the cases with a parseable baseline verdict.
Model
Part.
−1
−0.5
0
+0.5
+1
Llama-3.1-8B
LOW
92.3 ∗
92.5
92.8
92.5
92.3 ∗
Llama-3.1-8B
HIGH
79.7
82.2
82.2
82.7
82.2
Llama-3.3-70B
LOW
91.7 ∗∗
91.9
92.1
92.3
92.3
Llama-3.3-70B
HIGH
73.3 ∗
76.7
79.2
79.2
78.7
Table 3: Accuracy (%) by steering α , all four (model, partition) cells. Bold marks the unsteered baseline. ∗p<0.05 , ∗∗p<0.01 (McNemar’s exact test vs. the baseline; nominal, uncorrected for the 16 comparisons in this table). Exact p -values for every cell in Appendix H .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
n total
Modal cov.
gdpr
3,759
59.1%
hipaa
6,924
20.5%
eu_ai_act
7,758
60.7%
ccpa
2,489
30.9%
sb35
1,244
31.8%
finance_crypto
1,969
38.9%
Appendix
Table 4: Domain sizes and per-sample mandatory-modal coverage (OCS-neg applicability). The three domains at 0.0% (MITRE ATT&CK technique descriptions, academic-integrity value statements, and online-learning guidance) are not phrased in mandatory-obligation language at all. This is a genuine linguistic property of these domains, not a processing failure; with nothing to negate, they are excluded from OCS-neg entirely. † Domain name as in OmniCompliance-100K.
Model
Baseline
Del
Swap
Neg
Qwen2.5-7B
90.3%
90.7%
83.2%
79.5%
Llama-3.1-8B
91.9%
91.4%
89.4%
89.7%
Llama-3.3-70B
90.9%
87.5%
84.5%
83.6%
Llama-Guard-3-8B
51.1%
50.0%
50.1%
50.9%
Mistral-7B-v0.3
90.3%
91.6%
89.6%
88.7%
Appendix
Table 5: Baseline accuracy and accuracy under each rule perturbation condition, pooled across all 20 domains. Accuracy under rule deletion approximates performance when only case content is available.
Figure 3: Mean OCS by perturbation condition, pooled across all models and domains. OCS-del (rule deleted) is the lowest of the three: verdicts are least likely to change when the rule is removed entirely.
Model
mean ∣ ICS ∣
mean ∣ ICS-delta ∣
Ratio
Llama-3.1-8B
4.56
0.58
12.8%
Llama-3.3-70B
3.31
0.69
21.0%
Mistral-7B-v0.3
1.45
0.33
22.9%
Qwen2.5-7B
16.71
2.79
16.7%
Appendix
Table 6: Mean baseline ICS magnitude, mean ∣ ICS-delta ∣ , and their ratio, pooled across all three perturbation conditions and 20 domains.
Figure 4: ICS-delta as a percentage of baseline ICS magnitude, per model per condition. Every bar is a small fraction of the baseline signal; OCS-neg (rightmost, green) is consistently the smallest.
Model
ρ
p
Qwen2.5-7B-Instruct
0.114
0.631
Llama-3.1-8B-Instruct
0.084
0.724
Mistral-7B-Instruct-v0.3
−0.032
0.895
Llama-3.3-70B-Instruct
0.338
0.145
Appendix
Table 7: OCS-agg vs. ICS-delta-agg per-domain rank correlation, one row per Tier-1 model ( n=20 domains each). None reaches conventional significance.
Model
Null (indep.)
Obs. OCS
Perm. p
Llama-3.1-8B
0.500
0.070
<0.0001
Llama-3.3-70B
0.486
0.081
<0.0001
Mistral-7B-v0.3
0.502
0.089
<0.0001
Qwen2.5-7B
0.483
0.094
<0.0001
Llama-Guard-3-8B
0.019
0.013
<0.0001
Appendix
Table 8: Observed OCS-agg vs. an independence null built from each model’s own empirical baseline/altered verdict marginals ( p(1−q)+q(1−p) ), validated by a 10,000-sample within-domain permutation test ( p : one-sided, observed ≤ null). Every model’s OCS is significantly below its own bias-corrected null, not only below the naive 0.5 reference.
Model
Cond.
n
Obs.
Null
p↓
p↑
Llama-3.1-8B
swap
132
0.780
0.669
1.000
<0.0001
Llama-3.1-8B
neg
42
0.333
0.437
0.081
0.987
Llama-3.3-70B
swap
187
0.733
0.723
0.881
0.329
Llama-3.3-70B
neg
53
0.377
0.479
0.036
0.996
Mistral-7B-v0.3
swap
120
0.750
0.665
0.999
0.006
Mistral-7B-v0.3
neg
36
0.389
0.486
0.166
0.961
Appendix
Table 9: Full rule-necessity control statistics: OCS-swap and OCS-neg per model, restricted to that model’s own rule-necessary subset (baseline correct, del incorrect). Llama-3.3-70B-Instruct’s nominal p↓=0.036 for OCS-neg does not survive any correction for the ten tests conducted here (Bonferroni threshold 0.005 ); we do not treat it as a significant reduction.
Figure 5: OCS-agg vs. baseline accuracy, all 100 (model, domain) cells. The correlation is positive but narrowly misses p<0.05 .
Feature
Spearman ρ
pBH
Sig. (BH 0.05)
ECD
0.041
0.86
no
CRC
0.241
0.38
no
MMR
0.405
0.14
no
MDD
0.570
0.044
yes
FKGL
0.448
0.12
no
Appendix
Table 10: Linguistic feature correlations with per-domain mean OCS-agg ( n=19 –20 domains). Only MDD survives Benjamini-Hochberg correction.
Figure 6: Standardized OLS coefficients (all five features jointly, → OCS-agg). Red bars mark p<0.05 individually within the joint model; none reach this threshold, and MDD is closest.
Model
Δlow
Δhigh
p (LOW)
Llama-3.1-8B
−0.006
0.000
0.146
Llama-3.3-70B
+0.007
−0.033
0.073
Qwen2.5-7B
−0.028
−0.006
<0.0001
Mistral-7B-v0.3
−0.010
0.000
0.077
Llama-Guard-3-8B †
−0.001
0.000
0.500
Appendix
Table 11: Rule-grounded prompting: accuracy change vs. standard prompting on the LOW/HIGH-OCS partitions (Llama-3.3-70B-Instruct’s per-sample OCS), parseable-verdict cases only; p is McNemar’s exact test on LOW. Only Qwen2.5-7B reaches significance, and harmfully. † Robustness check, not a grounded-reasoning test (Section 4.3 ).
Model
Part.
α=−1
−0.5
+0.5
+1
Llama-3.1-8B
LOW
0.0489
0.1628
0.1496
0.0273
Llama-3.1-8B
HIGH
0.2668
1.0000
1.0000
1.0000
Llama-3.3-70B
LOW
0.0044
0.1094
0.1153
0.1742
Llama-3.3-70B
HIGH
0.0290
0.3323
1.0000
1.0000
Appendix
Table 12: Exact McNemar’s-test p -values, each (model, partition, α ) cell vs. that row’s α=0 baseline, paired per case.
Figure 7: Accuracy vs. steering α , LOW- and HIGH-OCS partitions, both models. Negative α degrades accuracy for both models on both partitions; positive α never significantly helps.