cs.AIOct 8, 2026

Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems

Authors: Saisab Sadhu, Aadit Sengupta, Vinay kumar Sankarapu, Pratinav Seth

Organizations: Lexsi Labs

Abstract

Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Code-as-Auditor: Executable Compliance Reasoning via Regulation-to-Code

    Sep 16, 2026Jisoo Kim, Taeyoon Kwack, Jinwoo Jang +2LLM GroundingLegal Reasoning

  2. FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions

    May 28, 2026Huaixia Dou, Jie Zhu, Minghao Wu +5Language Model Safety EvaluationLLM Guardrails