Who Holds the Pen? Let Specifications, Not Agents, Sign Off
Organizations: Department of Computer Science and Engineering, The University of Texas at Arlington, Arlington, TX, USA · Monash University, Melbourne, VIC, Australia · Department of Computer Science, Kent State University, Kent, OH, USA
Abstract
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an agent's interpretation or completion claim does not establish the required state. On SkillsBench, using only agent-visible prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions. Across seven models, only 79.6%--86.4% are satisfied, while completion-claim rates exceed official evaluator pass rates by 28.7--37.9 percentage points. We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this principle by compiling visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion.
Figures & tables
| Raw Agent | SpecHarness | Change vs. Raw (pp) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Pass | U–E | S–A | Pass | U–E | S–A | Pass | U–E | S–A |
| GPT-5.6 Sol | 71.3 | 13.6 | 28.7 | 85.1 | 6.3 | 6.9 | +13.8 | -7.3 | -21.8 |
| Claude Fable 5 | 69.0 | 14.5 | 31.0 | 81.6 | 7.1 | 10.3 | +12.6 | -7.4 | -20.7 |
| Gemini 3.1 Pro | 62.1 | 17.1 | 37.9 | 79.3 | 8.4 | 12.6 | +17.2 | -8.7 | -25.3 |
| Kimi K3 | 60.9 | 17.5 | 32.2 | 67.8 | 10.8 | 14.9 | +6.9 | -6.7 | -17.3 |
| GLM-5.2 | 57.5 | 18.7 | 29.9 | 69.0 | 9.6 | 11.5 | +11.5 | -9.1 | -18.4 |
| Method | Pass | U–E | S–A | Prov. | Eff. | Cmt. |
|---|---|---|---|---|---|---|
| Post-hoc Verification | ||||||
| Agentic Rubrics (ACL’26) | 74.7 | 12.6 | 23.0 | |||
| Completion Gating | ||||||
| VeriMAP (EACL’26) | 79.3 | 9.6 | 20.7 | |||
| Runtime Enforcement | ||||||
| AgentSpec (ICSE’26) | 78.2 | 8.8 | 24.1 | |||
| Raw SpecHarness | |||
|---|---|---|---|
| Model | Pass | U–E | S–A |
| GPT-5.6 Sol | 91.3 95.1 | 8.2 3.5 | 8.7 3.8 |
| Claude Fable 5 | 89.6 94.0 | 9.1 4.1 | 10.4 4.7 |
| Gemini 3.1 Pro | 88.2 93.2 | 8.7 3.8 | 11.8 5.2 |
| Kimi K3 | 85.7 90.6 | 11.6 5.9 | 14.3 7.4 |
| GLM-5.2 | 84.4 89.8 | 10.3 5.2 | 15.6 8.1 |
| Method | Pass | U–E | S–A | Prov. | Eff. | Cmt. |
|---|---|---|---|---|---|---|
| Post-hoc Verification | ||||||
| RvLLM (NeurIPS’25) | 90.5 | 7.0 | 7.8 | |||
| Completion Gating | ||||||
| VeriMAP (EACL’26) | 91.7 | 6.2 | 6.7 | |||
| Computation Substrate | ||||||
| SatLM (NeurIPS’23) | 92.3 | 4.9 | 6.4 | |||
| Variant | Pass | U–E | S–A | Pres. |
|---|---|---|---|---|
| Full SpecHarness | 85.1 | 6.3 | 6.9 | 96.8 |
| w/o Mediation | 80.5 | 10.0 | 11.5 | 95.2 |
| w/o Effect Validation | 78.2 | 9.4 | 16.1 | 93.5 |
| w/o Commitment | 81.6 | 8.1 | 25.3 | 98.4 |
| w/o Blocking Qualification | 79.3 | 8.8 | 5.7 | 90.3 |
| w/o Skill Obligations | 77.0 | 12.0 | 14.9 | 91.9 |
| Variant | Stale Acc. | Invalid. | Revalid. | Recovery |
|---|---|---|---|---|
| Full SpecHarness | 0.0 | 100.0 | 95.8 | 95.6 |
| w/o Freshness | 100.0 | 0.0 | N/A | N/A |
| Runtime Mode | SkillsBench | GuideBench |
|---|---|---|
| Mediate-and-commit | 44.0% | 0.0% |
| Validate-and-commit | 56.0% | 100.0% |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Compiler | Matched | Dev. Cov. | Accounted | Unit Map | Valid Rec. |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 77/81 | 95.1% | 98.8% | 100.0% | 97.6% |
| Claude Fable 5 | 75/81 | 92.6% | 100.0% | 98.5% | 98.7% |
| Gemini 3.1 | 73/81 | 90.1% | 97.4% | 96.9% | 96.1% |
| Kimi K3 | 71/81 | 87.7% | 98.6% | 95.4% | 97.2% |
| GLM-5.2 | 69/81 | 85.2% | 96.0% | 93.8% | 94.7% |
| Qwen3.7 | 67/81 | 82.7% | 97.2% | 92.3% | 96.0% |
| Audit Metric | GPT-5.6 Sol | Claude Fable 5 | Gemini 3.1 | Kimi K3 | GLM-5.2 | Qwen3.7 | DeepSeek-V4 |
|---|---|---|---|---|---|---|---|
| Directions | 509 | 521 | 496 | 517 | 481 | 468 | 492 |
| Direction-accounted tasks | 87/87 (100.00%) | 86/87 (98.85%) | 86/87 (98.85%) | 86/87 (98.85%) | 85/87 (97.70%) | 84/87 (96.55%) | 85/87 (97.70%) |
| Unit-mapped tasks | 87/87 (100.00%) | 85/87 (97.70%) | 86/87 (98.85%) | 85/87 (97.70%) | 84/87 (96.55%) | 83/87 (95.40%) | 84/87 (96.55%) |
| Broad test correspondence | 573/585 (97.95%) | 555/585 (94.87%) | 561/585 (95.90%) | 552/585 (94.36%) | 546/585 (93.33%) | 537/585 (91.79%) | 526/585 (89.91%) |
| Fine-grained subset | 406/585 (69.40%) | 382/585 (65.30%) | 389/585 (66.50%) | 374/585 (63.93%) | 365/585 (62.39%) | 349/585 (59.66%) | 337/585 (57.61%) |
| Below broad criterion | 12/585 (2.05%) | 30/585 (5.13%) | 24/585 (4.10%) | 33/585 (5.64%) | 39/585 (6.67%) | 48/585 (8.21%) | 59/585 (10.09%) |
| Validator Qualification | SkillsBench | GuideBench |
| Candidate validators | 278 | 104 |
| Qualified for blocking | 250 | 96 |
| Satisfying cases | 556 | 208 |
| Targeted violations | 834 | 312 |
| Malformed or missing cases | 278 | 104 |
| Provider or execution errors | 278 | 104 |
| Audit Class | Cases | Correct | Violations |
|---|---|---|---|
| Direct access | 220 | 220 | 0 |
| Alternate path | 176 | 174 | 2 |
| Path handling | 264 | 261 | 3 |
| Subprocess | 132 | 128 | 4 |
| Malformed proposal | 220 | 220 | 0 |
| Token scope | 220 | 220 | 0 |
| Runtime Qualification | Obligations | Share |
|---|---|---|
| Closure-audited mediation | 110/250 | 44.0% |
| Safely isolated validation | 140/250 | 56.0% |
| Total hard obligations | 250/250 | 100.0% |
| Mutation Class | Trials | Affected Entries | Invalidated | Revalidated | Affected Tasks | Recovered |
|---|---|---|---|---|---|---|
| Artifact content | 80 | 190 | 190 | 183 | 72 | 69 |
| Path or identity | 50 | 121 | 121 | 116 | 46 | 44 |
| Schema or configuration | 46 | 108 | 108 | 103 | 42 | 40 |
| Upstream obligation | 42 | 104 | 104 | 99 | 39 | 37 |
| Validator or environment | 30 | 72 | 72 | 69 | 29 | 28 |
| Total | 248 | 595 | 595 | 570 | 228 | 218 |
| Agent | Provider Model ID |
|---|---|
| GPT-5.6 Sol | gpt-5.6-sol |
| Claude Fable 5 | claude-fable-5 |
| Gemini 3.1 Pro | gemini-3.1-pro-preview |
| Kimi K3 | kimi-k3 |
| GLM-5.2 | glm-5.2 |
| Qwen3.7-Max | qwen3.7-max |
| Benchmark | Condition | Runtime Evidence | Validator Feedback | Action Mediation | State Commitment | Completion Rule |
|---|---|---|---|---|---|---|
| SkillsBench | Raw | No | No | No | No | Ungated agent claim |
| SkillsBench | Agentic Rubrics | No | No | No | No | Ungated agent claim |
| SkillsBench | VeriMAP | Terminal | Gate result | No | No | Verifier gate |
| SkillsBench | AgentSpec | Policy state | Rule decision | Yes | No | Native termination |
| SkillsBench | SpecHarness | Online | Source-linked | Closure-audited | Versioned | Fresh mandatory satisfaction |
| GuideBench | Raw | No | No | N/A | No | Ungated agent claim |
| Benchmark | Metric | Change | 95% CI |
|---|---|---|---|
| SkillsBench | Pass | ||
| U–E | |||
| S–A | |||
| GuideBench | Pass | ||
| U–E | |||
| S–A |
| Agent | Raw Only | SpecHarness Only | Exact | Holm |
|---|---|---|---|---|
| GPT-5.6 Sol | 1 | 13 | 0.00183 | 0.01099 |
| Claude Fable 5 | 1 | 12 | 0.00342 | 0.01709 |
| Gemini 3.1 Pro | 1 | 16 | 0.00027 | 0.00192 |
| Kimi K3 | 0 | 6 | 0.03125 | 0.03125 |
| GLM-5.2 | 1 | 11 | 0.00635 | 0.02539 |
| Qwen3.7-Max | 1 | 11 | 0.00635 | 0.02539 |
| SkillsBench Category | Units | Raw U–E | Spec. U–E | Change |
|---|---|---|---|---|
| Artifact content | 80 | 17.0 | 8.4 | |
| Structural constraints | 105 | 18.5 | 9.7 | |
| Numeric constraints | 92 | 16.8 | 8.6 | |
| Preservation | 74 | 15.9 | 8.2 | |
| Behavioral requirements | 86 | 18.8 | 10.6 | |
| Order and coverage | 72 | 16.9 | 9.9 |
| GuideBench Category | Instances | Raw U–E | Spec. U–E | Change |
|---|---|---|---|---|
| Applicability | 1,300 | 10.1 | 4.8 | |
| Priority | 1,100 | 10.7 | 5.0 | |
| Exceptions | 900 | 11.2 | 5.5 | |
| Required response | 1,200 | 10.3 | 4.9 | |
| Output constraints | 800 | 9.8 | 4.7 | |
| Other rule relations | 517 | 10.4 | 5.1 |
| SkillsBench Residual Failure | Count |
|---|---|
| Outside hard-enforcement scope | 35 |
| Qualified obligation unresolved | 28 |
| Validator error or unavailable | 22 |
| Repair or interaction exhausted | 51 |
| Held-out evaluator mismatch | 28 |
| Total | 164 |
| GuideBench Residual Failure | Count |
|---|---|
| Outside hard-enforcement scope | 125 |
| Qualified obligation unresolved | 96 |
| Validator error or unavailable | 173 |
| Repair or interaction exhausted | 147 |
| Held-out evaluator mismatch | 96 |
| Total | 637 |
| Agent | Raw Tokens/Task | SpecHarness Tokens/Task | Token Ratio | Time Ratio | Calls/Task Difference |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 68,161 | 104,286 | 1.53 | 1.24 | |
| Claude Fable 5 | 65,000 | 97,500 | 1.50 | 1.25 | |
| Gemini 3.1 Pro | 72,000 | 113,040 | 1.57 | 1.22 | |
| Kimi K3 | 69,000 | 106,950 | 1.55 | 1.26 | |
| GLM-5.2 | 64,000 | 97,280 | 1.52 | 1.23 | |
| Qwen3.7-Max | 70,000 | 107,800 | 1.54 | 1.25 |
| Diagnostic Measure | Raw | SpecHarness |
|---|---|---|
| Actionable source-linked reports (%) | 42.5 | 96.2 |
| Normalized feedback time | 1.00 | 0.44 |
| Benchmark | Visible Specification | Governed State | Qualified Evidence | Runtime Mode |
|---|---|---|---|---|
| SkillsBench | Task, workspace, and skill | Artifact, execution, and completion state | Controlled execution and workspace validators | Mediate-and-commit or validate-and-commit |
| GuideBench | Guideline and decision context | Rule-local decision and completion state | Rule-application and decision validators | Validate-and-commit |
| Step | Runtime Boundary | SkillsBench: data-to-d3 | GuideBench: chat[176] | Authority Invariant |
|---|---|---|---|---|
| 1 | Source compilation | Compile the required output paths and local D3 dependency from task.md and the visible D3 skill. | Compile rule_6 : an existing booking plus cancellation intent requires cancellation and an explicit success confirmation. | Every obligation retains an agent-visible source anchor. |
| 2 | Disposition and qualification | Mark mandatory and bind it to , whose inputs are observable workspace paths. | Mark mandatory and bind its applicability and decision checks to the same rule-local obligation. | Only grounded requirements with qualified evidence providers enter hard enforcement. |
| 3 | Dependency binding | Bind to index.html , the local JS and CSS files, and copied data at workspace version . | Bind to rule_6 , the dialogue context, and the candidate-option set at input version . | Evidence remains valid only for its recorded dependency versions. |
| 4 | Agent proposal | Propose controlled writes under /root/output . | Propose option A: the dialogue follows the cancellation guideline. | Agent intent cannot directly update authoritative state. |
| 5 | Obligation matching | Match the proposed writes to and return for the canonical output paths. | Activate ; rules whose antecedents are false remain inactive. | Authorization and applicability are scoped to matched obligations. |
| 6 | Controlled execution | Execute the allowed writes through the closure-audited file surface; writes outside the authorized path set require a new decision. | No physical action is mediated; the proposed answer remains a validate-and-commit decision. | Agent intent cannot bypass the applicable runtime procedure. |