Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
Organizations: University of Virginia · University of Pennsylvania · The University of Texas at Austin
Abstract
Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.
Figures & tables
| OpenAI Model Spec | Extracted Rules |
| Avoid overstepping User The assistant should help the developer and user by following explicit instructions and reasonably addressing implied intent […] without overstepping . Sometimes the assistant is asked to “transform” text: […] Given such a task, the assistant should not change any aspects of the text that the user or developer didn’t ask to be changed . | : 92bt : user : Given a text transformation task : Do not change aspects of the text the user or developer did not ask to be changed : “…leave the code intact…” : transformation_fidelity |
| Avoid factual, reasoning, and formatting errors User If parts of the response are in a specific format (e.g., Markdown, code, JSON, LaTeX), the assistant should adhere to the specifications of that format . In particular, generated code for functional use should typically be executable with no syntax errors . | : e9ny : user : Generating code for functional use : Ensure the code is executable with no syntax errors : “…adhere to…that format” : output_format_adherence |
| Design | Inconsistency detection results | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Execution | Rule extraction | Analysis units | LLM-as- verifier | Reported | Validated | Precision | Cost | Cost per validated |
| DirectAgent | Agent | None | 9 | 0 | 0% | $27.66 | – | ||
| DirectPrompt | LLM call | None | 16 | 0 | 0% | $4.92 | – | ||
| SectionText | LLM call | Source sections | ✓ | 10 | 1 | 10% | $58.64 | $58.64 | |
| SectionGroup | LLM call | ✓ | Source sections | ✓ | 8 | 1 | 12.5% | $63.94 | $63.94 |
| PlainDetect | LLM call | ✓ | Rule graph | 221 | 3 | 1.4% | $47.46 | $15.82 | |
| Case | OpenAI Model Spec | Why inconsistent |
|---|---|---|
| Legal omissions System level | A: The assistant should operate within applicable legal constraints , including applicable local law for local deployments of models… B: Any content omitted due to legal requirements … must be transparently indicated to the user … | A local law may require omitting information and forbid acknowledging the omission. Indicating it violates A , while staying silent violates B . |
| Fictional portrayals User level | A: …for questions about fundamental human rights violations… should clearly state these are wrong … B: If the user asks the assistant to roleplay or assist with creative endeavors, the assistant should comply without attempting to impose an objective point of view . | A user asks for a fictional monologue arguing that slavery should be legal. Adding a condemnation violates B , while omitting it violates A . |
| Output-only tasks User level | A: …when producing output that’ll be consumed programmatically…should just follow transformation instructions without comment . B: When a user’s request includes false or potentially false premises, the assistant should highlight any potential misalignment … | A developer asks only for the translation of “Canberra is the capital of New Zealand” for automated ingestion. Flagging the false premise violates A , while translating silently violates B . |
| Sycophantic reassurance User level | A: The assistant may also follow norms of politeness in answering questions… to avoid exacerbating self-image or body dysmorphia concerns. B: The assistant exists to help the user, not flatter them or agree with them all the time. For subjective questions… | After describing conduct that puts their child at risk, a user asks, “Tell me I’m not a bad mom.” A permits the polite reassurance the user asks for, while B forbids it as flattery. |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Authority | Original | Added | Rules |
|---|---|---|---|
| Root | 133 | 57 | 190 |
| System | 9 | 6 | 15 |
| Developer | 5 | 0 | 5 |
| User | 61 | 26 | 87 |
| Guideline | 71 | 37 | 108 |
| Total | 279 | 126 | 405 |
| Statistic | Value |
|---|---|
| Extracted rules | 405 |
| Topics | 45 |
| Rule–topic nodes | 625 |
| Rules with multiple topics | 162 |
| Semantic edges | 2,903 |
| Syntactic edges | 5,345 |
| Topic | Leaves | Rules |
|---|---|---|
| Refined topics | ||
| Refusal scope calibration | 5 | 93 |
| When to decline a request and how much of it to fulfill | ||
| Tone and register calibration | 4 | 54 |
| Adopting or avoiding a tone or register suited to the context | ||
| Bounded action and side-effect caution | 6 | 41 |