Who Holds the Pen? Let Specifications, Not Agents, Sign Off
Authors: Haiqing Li, Xin Ma, Yinhao Wu, Wenliang Zhong, Feng Jiang, Thao M. Dang, Xiao Hu, Hehuan Ma, +2 more
Organizations: Department of Computer Science and Engineering, The University of Texas at Arlington, Arlington, TX, USA · Monash University, Melbourne, VIC, Australia · Department of Computer Science, Kent State University, Kent, OH, USA
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an agent's interpretation or completion claim does not establish the required state. On SkillsBench, using only agent-visible prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions. Across seven models, only 79.6%--86.4% are satisfied, while completion-claim rates exceed official evaluator pass rates by 28.7--37.9 percentage points. We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this principle by compiling visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion.
Figures & tables
Figure 1: Two structural gaps in specification following. Top: The agent understands the required order but executes the steps incorrectly, illustrating the understanding–execution gap. Bottom: The agent claims completion before the specification-required state is established, illustrating the state–authority gap.
Figure 2: Three intervention paradigms for specification compliance. (a) Post-hoc verification detects violations only after execution has completed. (b) Completion gating checks the agent’s completion claim before acceptance but does not constrain the preceding execution. (c) Runtime enforcement intervenes during execution to prevent or constrain specification-violating actions.
Figure 3: Empirical evidence for the two structural gaps across seven language models. Top: Models satisfy 79.6%–86.4% of the 509 source-grounded task directions recovered from agent-visible materials. Bottom: Agent-reported completion rates exceed official verifier pass rates by 28.7–37.9 percentage points.
Figure 4: SpecHarness compiles agent-visible specifications into source-linked obligations, mediates closure-audited actions, validates observed effects, commits versioned state from authorized evidence, and permits finalization only when all fresh mandatory obligations are satisfied. Advisory and abstained requirements remain outside hard enforcement.
Raw Agent
SpecHarness
Change vs. Raw (pp)
Model
Pass ↑
U–E ↓
S–A ↓
Pass ↑
U–E ↓
S–A ↓
Pass
U–E
S–A
GPT-5.6 Sol
71.3
13.6
28.7
85.1
6.3
6.9
+13.8
-7.3
-21.8
Claude Fable 5
69.0
14.5
31.0
81.6
7.1
10.3
+12.6
-7.4
-20.7
Gemini 3.1 Pro
62.1
17.1
37.9
79.3
8.4
12.6
+17.2
-8.7
-25.3
Kimi K3
60.9
17.5
32.2
67.8
10.8
14.9
+6.9
-6.7
-17.3
GLM-5.2
57.5
18.7
29.9
69.0
9.6
11.5
+11.5
-9.1
-18.4
Table 1: Cross-model results on all 87 SkillsBench tasks. Pass is the official-verifier pass rate; U–E and S–A are the understanding–execution and state–authority gaps. Change columns report percentage-point differences from Raw. Macro Average is the unweighted mean across seven task-agent models.
Method
Pass ↑
U–E ↓
S–A ↓
Prov.
Eff.
Cmt.
Post-hoc Verification
Agentic Rubrics (ACL’26)
74.7
12.6
23.0
△
△
×
Completion Gating
VeriMAP (EACL’26)
79.3
9.6
20.7
△
✓
×
Runtime Enforcement
AgentSpec (ICSE’26)
78.2
8.8
24.1
✓
×
×
Table 2: SkillsBench paradigm comparison with GPT-5.6 Sol fixed. Pass is the official-verifier pass rate; U–E and S–A are the two structural gaps. Prov., Eff., and Cmt. denote source-linked provenance, observable-effect validation, and versioned authoritative commitment. Markers indicate full ( ✓ ), partial or adaptation-dependent ( △ ), or no support ( × ) in the evaluated adaptations.
Raw → SpecHarness
Model
Pass ↑
U–E ↓
S–A ↓
GPT-5.6 Sol
91.3 → 95.1
8.2 → 3.5
8.7 → 3.8
Claude Fable 5
89.6 → 94.0
9.1 → 4.1
10.4 → 4.7
Gemini 3.1 Pro
88.2 → 93.2
8.7 → 3.8
11.8 → 5.2
Kimi K3
85.7 → 90.6
11.6 → 5.9
14.3 → 7.4
GLM-5.2
84.4 → 89.8
10.3 → 5.2
15.6 → 8.1
Table 3: GuideBench results across seven agents. Each cell reports Raw → SpecHarness; U–E and S–A denote the two decision-state gaps.
Method
Pass ↑
U–E ↓
S–A ↓
Prov.
Eff.
Cmt.
Post-hoc Verification
RvLLM (NeurIPS’25)
90.5
7.0
7.8
△
△
×
Completion Gating
VeriMAP (EACL’26)
91.7
6.2
6.7
△
✓
×
Computation Substrate
SatLM (NeurIPS’23)
92.3
4.9
6.4
×
△
×
Table 4: GuideBench paradigm comparison with GPT-5.6 Sol fixed. Metrics and markers follow Table 2 .
Variant
Pass ↑
U–E ↓
S–A ↓
Pres. ↑
Full SpecHarness
85.1
6.3
6.9
96.8
w/o Mediation
80.5
10.0
11.5
95.2
w/o Effect Validation
78.2
9.4
16.1
93.5
w/o Commitment
81.6
8.1
25.3
98.4
w/o Blocking Qualification
79.3
8.8
5.7
90.3
w/o Skill Obligations
77.0
12.0
14.9
91.9
Table 5: Runtime architecture ablation on SkillsBench with GPT-5.6 Sol fixed. Pres. denotes paired Raw-pass preservation.
Variant
Stale Acc. ↓
Invalid. ↑
Revalid. ↑
Recovery ↑
Full SpecHarness
0.0
100.0
95.8
95.6
w/o Freshness
100.0
0.0
N/A
N/A
Table 6: Freshness invalidation and recovery (%) over 248 paired targeted SkillsBench mutations; protocol and denominators appear in Appendix B.3.
Runtime Mode
SkillsBench
GuideBench
Mediate-and-commit
44.0%
0.0%
Validate-and-commit
56.0%
100.0%
Table 7: Runtime-mode allocation among constructed hard obligations; GuideBench uses validate-and-commit exclusively.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Compiler
Matched
Dev. Cov. ↑
Accounted ↑
Unit Map ↑
Valid Rec. ↑
GPT-5.6 Sol
77/81
95.1%
98.8%
100.0%
97.6%
Claude Fable 5
75/81
92.6%
100.0%
98.5%
98.7%
Gemini 3.1
73/81
90.1%
97.4%
96.9%
96.1%
Kimi K3
71/81
87.7%
98.6%
95.4%
97.2%
GLM-5.2
69/81
85.2%
96.0%
93.8%
94.7%
Qwen3.7
67/81
82.7%
97.2%
92.3%
96.0%
Appendix
Table 8: Candidate-compiler results on 14 task-disjoint development tasks with 65 source units and 81 annotated directions. Matched is the numerator of development coverage. Accounted and Valid Rec. use extracted records as their denominator; Unit Map uses the 65 source units.
Audit Metric
GPT-5.6 Sol
Claude Fable 5
Gemini 3.1
Kimi K3
GLM-5.2
Qwen3.7
DeepSeek-V4
Directions
509
521
496
517
481
468
492
Direction-accounted tasks
87/87 (100.00%)
86/87 (98.85%)
86/87 (98.85%)
86/87 (98.85%)
85/87 (97.70%)
84/87 (96.55%)
85/87 (97.70%)
Unit-mapped tasks
87/87 (100.00%)
85/87 (97.70%)
86/87 (98.85%)
85/87 (97.70%)
84/87 (96.55%)
83/87 (95.40%)
84/87 (96.55%)
Broad test correspondence
573/585 (97.95%)
555/585 (94.87%)
561/585 (95.90%)
552/585 (94.36%)
546/585 (93.33%)
537/585 (91.79%)
526/585 (89.91%)
Fine-grained subset
406/585 (69.40%)
382/585 (65.30%)
389/585 (66.50%)
374/585 (63.93%)
365/585 (62.39%)
349/585 (59.66%)
337/585 (57.61%)
Below broad criterion
12/585 (2.05%)
30/585 (5.13%)
24/585 (4.10%)
33/585 (5.64%)
39/585 (6.67%)
48/585 (8.21%)
59/585 (10.09%)
Appendix
Table 9: Post-selection construction and official-test correspondence for the seven frozen candidates. Task-level rates use 87 tasks; correspondence rates use 585 official test functions. These results do not participate in compiler selection.
Validator Qualification
SkillsBench
GuideBench
Candidate validators
278
104
Qualified for blocking
250
96
Satisfying cases
556
208
Targeted violations
834
312
Malformed or missing cases
278
104
Provider or execution errors
278
104
Appendix
Table 10: Validator qualification before task-agent evaluation. Expected outcomes count cases producing the prespecified result.
Audit Class
Cases
Correct
Violations
Direct access
220
220
0
Alternate path
176
174
2
Path handling
264
261
3
Subprocess
132
128
4
Malformed proposal
220
220
0
Token scope
220
220
0
Appendix
Table 11: Closure stress tests over candidate mediated surfaces. Detected violations exclude the corresponding paths or surfaces from closure-qualified mediation.
Runtime Qualification
Obligations
Share
Closure-audited mediation
110/250
44.0%
Safely isolated validation
140/250
56.0%
Total hard obligations
250/250
100.0%
Appendix
Table 12: Runtime-mode qualification of constructed SkillsBench hard obligations.
Mutation Class
Trials
Affected Entries
Invalidated
Revalidated
Affected Tasks
Recovered
Artifact content
80
190
190
183
72
69
Path or identity
50
121
121
116
46
44
Schema or configuration
46
108
108
103
42
40
Upstream obligation
42
104
104
99
39
37
Validator or environment
30
72
72
69
29
28
Total
248
595
595
570
228
218
Appendix
Table 13: Freshness results by dependency-mutation class. All 595 affected entries are invalidated; 570/595 (95.8%) are revalidated, and 218/228 (95.6%) affected tasks recover after repair.
Agent
Provider Model ID
GPT-5.6 Sol
gpt-5.6-sol
Claude Fable 5
claude-fable-5
Gemini 3.1 Pro
gemini-3.1-pro-preview
Kimi K3
kimi-k3
GLM-5.2
glm-5.2
Qwen3.7-Max
qwen3.7-max
Appendix
Table 14: Provider model identifiers used for the evaluated task agents. The same model endpoint is retained within each paired comparison.
Benchmark
Condition
Runtime Evidence
Validator Feedback
Action Mediation
State Commitment
Completion Rule
SkillsBench
Raw
No
No
No
No
Ungated agent claim
SkillsBench
Agentic Rubrics
No
No
No
No
Ungated agent claim
SkillsBench
VeriMAP
Terminal
Gate result
No
No
Verifier gate
SkillsBench
AgentSpec
Policy state
Rule decision
Yes
No
Native termination
SkillsBench
SpecHarness
Online
Source-linked
Closure-audited
Versioned
Fresh mandatory satisfaction
GuideBench
Raw
No
No
N/A
No
Ungated agent claim
Appendix
Table 15: Runtime access and completion semantics of the evaluated adaptations. Runtime Evidence denotes evidence available to the condition during execution; all conditions are evaluated afterward by the same frozen measurement validators. Entries describe the implemented conditions rather than every capability of the original systems.
Benchmark
Metric
Change
95% CI
SkillsBench
Pass
+12.0
[+9.1,+14.9]
U–E
−8.1
[−9.0,−7.2]
S–A
−20.0
[−22.8,−17.4]
GuideBench
Pass
+5.1
[+4.4,+5.8]
U–E
−5.4
[−6.0,−4.8]
S–A
−6.9
[−7.6,−6.2]
Appendix
Table 16: Paired task-bootstrap changes from Raw to SpecHarness in percentage points. Intervals use 10,000 task-level replicates.
Agent
Raw Only
SpecHarness Only
Exact p
Holm p
GPT-5.6 Sol
1
13
0.00183
0.01099
Claude Fable 5
1
12
0.00342
0.01709
Gemini 3.1 Pro
1
16
0.00027
0.00192
Kimi K3
0
6
0.03125
0.03125
GLM-5.2
1
11
0.00635
0.02539
Qwen3.7-Max
1
11
0.00635
0.02539
Appendix
Table 17: Exact paired McNemar tests for SkillsBench Pass. Raw Only and SpecHarness Only are discordant task counts; p -values are two-sided.
SkillsBench Category
Units
Raw U–E
Spec. U–E
Change
Artifact content
80
17.0
8.4
−8.6
Structural constraints
105
18.5
9.7
−8.8
Numeric constraints
92
16.8
8.6
−8.2
Preservation
74
15.9
8.2
−7.7
Behavioral requirements
86
18.8
10.6
−8.2
Order and coverage
72
16.9
9.9
−7.0
Appendix
Table 18: SkillsBench U–E by requirement category (%).
GuideBench Category
Instances
Raw U–E
Spec. U–E
Change
Applicability
1,300
10.1
4.8
−5.3
Priority
1,100
10.7
5.0
−5.7
Exceptions
900
11.2
5.5
−5.7
Required response
1,200
10.3
4.9
−5.4
Output constraints
800
9.8
4.7
−5.1
Other rule relations
517
10.4
5.1
−5.3
Appendix
Table 19: GuideBench U–E by rule-local analysis category (%).
SkillsBench Residual Failure
Count
Outside hard-enforcement scope
35
Qualified obligation unresolved
28
Validator error or unavailable
22
Repair or interaction exhausted
51
Held-out evaluator mismatch
28
Total
164
Appendix
Table 20: Residual SkillsBench official failures under SpecHarness, aggregated over seven models. The total equals the sum of the per-model official-failure counts.
GuideBench Residual Failure
Count
Outside hard-enforcement scope
125
Qualified obligation unresolved
96
Validator error or unavailable
173
Repair or interaction exhausted
147
Held-out evaluator mismatch
96
Total
637
Appendix
Table 21: Residual GuideBench official failures under SpecHarness, aggregated over seven models. The total satisfies 637=7×1,042−6,657 , where 6,657 is the aggregate official-pass count.
Agent
Raw Tokens/Task
SpecHarness Tokens/Task
Token Ratio
Time Ratio
Calls/Task Difference
GPT-5.6 Sol
68,161
104,286
1.53 ×
1.24 ×
+0.36
Claude Fable 5
65,000
97,500
1.50 ×
1.25 ×
+0.34
Gemini 3.1 Pro
72,000
113,040
1.57 ×
1.22 ×
+0.40
Kimi K3
69,000
106,950
1.55 ×
1.26 ×
+0.38
GLM-5.2
64,000
97,280
1.52 ×
1.23 ×
+0.35
Qwen3.7-Max
70,000
107,800
1.54 ×
1.25 ×
+0.37
Appendix
Table 22: Condition-execution cost over the 87 SkillsBench tasks. Ratios and call differences are computed separately for each task-agent model and then macro-averaged without model weighting.
Diagnostic Measure
Raw
SpecHarness
Actionable source-linked reports (%)
42.5
96.2
Normalized feedback time
1.00
0.44
Appendix
Table 23: Diagnostic information exposed by the complete conditions. The comparison does not assume matched trace access.
Benchmark
Visible Specification
Governed State
Qualified Evidence
Runtime Mode
SkillsBench
Task, workspace, and skill
Artifact, execution, and completion state
Controlled execution and workspace validators
Mediate-and-commit or validate-and-commit
GuideBench
Guideline and decision context
Rule-local decision and completion state
Rule-application and decision validators
Validate-and-commit
Appendix
Table 24: Instantiation of the SpecHarness authority model in SkillsBench and GuideBench.
Step
Runtime Boundary
SkillsBench: data-to-d3
GuideBench: chat[176]
Authority Invariant
1
Source compilation
Compile the required output paths and local D3 dependency from task.md and the visible D3 skill.
Compile rule_6 : an existing booking plus cancellation intent requires cancellation and an explicit success confirmation.
Every obligation retains an agent-visible source anchor.
2
Disposition and qualification
Mark oS mandatory and bind it to qS , whose inputs are observable workspace paths.
Mark oG mandatory and bind its applicability and decision checks to the same rule-local obligation.
Only grounded requirements with qualified evidence providers enter hard enforcement.
3
Dependency binding
Bind oS to index.html , the local JS and CSS files, and copied data at workspace version v0 .
Bind oG to rule_6 , the dialogue context, and the candidate-option set at input version g0 .
Evidence remains valid only for its recorded dependency versions.
4
Agent proposal
Propose controlled writes under /root/output .
Propose option A: the dialogue follows the cancellation guideline.
Match the proposed writes to oS and return u1=\textscallow for the canonical output paths.
Activate oG ; rules whose antecedents are false remain inactive.
Authorization and applicability are scoped to matched obligations.
6
Controlled execution
Execute the allowed writes through the closure-audited file surface; writes outside the authorized path set require a new decision.
No physical action is mediated; the proposed answer remains a validate-and-commit decision.
Agent intent cannot bypass the applicable runtime procedure.
Appendix
Table 25: Representative obligation-level protocol replays embedded in the end-to-end SkillsBench and GuideBench workflows. Each row illustrates an authority boundary for the selected obligation; other active obligations are omitted for space. A committed passed result contributes authoritative evidence, while satisfaction additionally requires the obligation predicate and current freshness. Finalization requires satisfaction of all active mandatory obligations. The examples use actual agent-visible benchmark materials, and the SkillsBench dependency mutation is controlled.
Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to read that text, or move the state into a module that enforces it, and each system is evaluated as a whole, so no one knows how much reliability comes from the state being shown, told, or enforced. We fix the task rules, the model, and paired episodes and vary how strongly task state reaches the agent: a raw transcript, an exact checklist, per-turn directives from a state machine compiled from the brief and advanced only by execution receipts, or an enforcement gate on that machine that refuses state-violating actions; every episode is scored by exact payload matching against dynamic ground truth. Across three models, two reasoning regimes, and two domains, four findings hold without per-turn reasoning: displaying accurate state is unreliable, an unverified ledger the agent writes itself beats an accurate checklist it is shown, directives help in proportion to the model's obedience, and enforcement needs no obedience but is bounded by the correctness of its state and by the matcher that maps requests to steps; per-turn reasoning at a 235B agent compresses these separations without repairing the text rungs. The same gate, compiled from τ2-bench's airline policy, raises a 235B agent's pass1 from 0.39 to 0.54 and changes nothing for a 35B agent that rarely violates the policy; on PM-Bench, where acting turns on recognizing a cue rather than on state, showing the record is the best rung--matching or beating both gates and reversing the ledger-over-checklist finding--and enforcing the matcher's judgement drops a 35B agent below its raw transcript. Enforcement pays when failures are state-decidable and frequent, and hurts when the gate's judgement is wrong.
Chenyu Zhang, Wonbin Kweon, Jiawei Han
University of Waterloo · Sungkyunkwan University · University of Illinois Urbana-Champaign
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document constrains its behavior over an extended tool-use horizon. We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks. Each task places an agent in a self-contained company environment (a file workspace with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol) and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20-124 pages. Tasks span five domains (finance, medical billing, insurance, logistics, and HR) and 10 fictional companies. To resist memorization, every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies. Grading is fully deterministic: each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not. Under strict grading, where a trial passes only if every criterion is satisfied, the strongest evaluated model passes 36.2% of trials, and most frontier models remain below 25%. Failures follow consistent patterns: agents let a plausible but unauthorized in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve. We release the tasks, environments, and evaluation harness.
Liudas Panavas, Sebastian Minus, Bradley Monton +4
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Caiqi Zhang, Rujun Han, Zifeng Wang +4
University of Cambridge · Google Cloud AI Research