Organizations: Newcastle University, UK · University of Bristol, UK · CSIRO, Australia · University College London, UK · Ohio State University, USA · University of Oxford, UK · FLock.io, UK · The University of Manchester, UK
Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides little evidence that the delivered work actually satisfied the task. We systematize this problem across the full lifecycle of an agent task. Our study organizes security and economic requirements into 17 property families over six stages, with receipt soundness and completeness assessed separately. We examine 12 systems and standards, five reusable mechanism families, and four classical baselines. We introduce guarantee closure, a task-relative criterion for determining whether guarantees established at one stage remain available and constrain the later decisions that depend on them. We apply the criterion to controlled and native workflows, covering 840 matched executions and an exhaustive 11,648-case check over a finite objective-task domain. Our results expose recurring failures between verification and settlement, where conforming work can remain unaccepted or valid evidence can be ignored. Public records and model judgments further distinguish recorded approval from evidence of task conformance, while economic analysis identifies the report, penalty, and shared-error assumptions behind these guarantees. These findings show where end-to-end guarantees fail and what must be repaired to preserve them across the workflow.
Figures & tables
Fig. 1: \tl_set:Ne dae Task lifecycle : agent roles, mechanisms, security and economic requirements, and threats across six stages. Negotiation fixes the acceptance predicate, verification produces a receipt binding the task, output, predicate, and supporting evidence, and settlement must validate this receipt before releasing payment. The annotated gaps illustrate how locally correct mechanisms can leave required guarantees missing, ignored, or unestablished.
Stage
P
Property
O,1D
O,XD
S,1D
S,XD
S1
P1
Sybil resistance
C
C
C
C
P2
Capability soundness
C
C
C
C
P3
Unforgeable, portable reputation
C
C
C
C
S2
P4
Incentive compatibility
C
C
C
C
P5
Individual rationality
⚫
⚫
⚫
⚫
P6
Weak budget balance
⚫
⚫
⚫
⚫
TABLE I: Stage-indexed requirements by task class.
Comparison unit
Property
F
R
U
-
Systems/standards (12)
P9a
0
7
5
0
Systems/standards (12)
P9b
0
4
4
4
Systems/standards (12)
P16
0
8
4
0
Mechanism families (5)
P9a
1
3
1
0
Mechanism families (5)
P9b
0
4
0
1
Mechanism families (5)
P16
0
5
0
0
TABLE II: Support counts by comparison unit.
Group
Artifact / Mechanism family
Class
P1 Sybil
P2 Capability
P3 Reputation
P4 IC
P5 IR
P6 BB
P7 Collusion
P8 Persuasion
P9a Sound
P9b Complete
P10 Confid.
P11 Metering
P12 Atomicity
P13 Timeliness
P14 Ordering
P15 Account.
P16 Evidence
P17 Capture
S1 Discovery
S2 Negotiation
S3 and S4 Exec./Verify
S5 Settlement
S6 Dispute
Systems and standards (12)
ERC-8004 draft [ 10 ]
Identity/reputation registry
❍
❍
◐
-
-
-
❍
-
◐
-
-
-
-
-
-
◐
◐
❍
ACP / Virtuals [ 4 ]
Escrow + evaluator
❍
❍
◐
❍
◐
⚫
❍
❍
◐
❍
❍
-
-
◐
❍
◐
◐
❍
ERC-8183 draft [ 5 ]
Job state machine
-
-
◐
❍
◐
⚫
❍
❍
◐
❍
◐
-
-
◐
◐
◐
◐
❍
AP2 v0.2 [ 7 ]
Verifiable payment credentials
-
-
-
-
◐
⚫
❍
❍
❍
-
◐
-
-
◐
◐
◐
◐
-
x402 v2 [ 6 , 30 , 31 , 32 ]
Web posted-price payments
-
❍
-
-
◐
⚫
❍
❍
❍
-
-
◐
-
◐
◐
◐
◐
-
TABLE III: Mechanism support for the lifecycle properties.
TABLE VI: Verification conditions for P9a and P9b.
Assessed unit
P9a
P9b
P16
Decisive evidence and retained condition
ERC-8004 registry
R
-
R
Validator responses and attributed feedback support provenance. Task conformance depends on the validator and linked evidence. The registry does not offer an acceptance service [ 10 ] .
ACP task protocol
R
U
R
Agreement, evaluation, and payment records bind a task history. Quality remains evaluator-dependent, and the documented evaluation route provides no bounded alternative when the evaluator withholds approval [ 4 ] .
ERC-8183 job escrow
R
U
R
Submission and evaluator-controlled terminal transitions attribute the verdict. The evaluator supplies quality evidence, while expiry returns client funds. Provider access to acceptance remains unestablished [ 5 ] .
AP2 payment mandates
U
-
R
Signed mandates establish payment authority and part of the decision history. The assessed interface supplies neither a conformance test nor an acceptance service [ 7 ] .
x402 payment and receipt
U
-
R
Signed offers and transfer or delivery statements bind payment terms. Task-conformance evidence and an acceptance service are separate workflow requirements [ 30 , 32 ] .
A402 attested exchange
R
R
R
Attested computation and request-bound exchange support a receipt under hardware trust. Exchange progress retains eventual-delivery and resource assumptions, while task-specific receipt deadlines and total cost bounds require separate support [ 33 ] .
Appendix
TABLE VII: EVIDENCE FOR ACCEPTANCE AND DECISION-RECORD RATINGS
Group
Work
Type
S1 Discovery
S2 Negotiation/terms
S3 Execution
S4 Verification
S5 Settlement
S6 Dispute/accountability
Formalized properties
Requirement derivation
System–property matrix
Cross-stage dependencies
End-to-end preservation
Strategic incentives
LLM robustness
Empirical validation
Task lifecycle
Property analysis
Composition
Agent/econ.
Evidence
Threats (3)
Mao et al. [ 15 ]
SoK
◐
◐
◐
◐
⚫
◐
❍
⚫
⚫
⚫
◐
⚫
⚫
❍
Shi et al. [ 109 ]
SoK
◐
❍
⚫
❍
❍
❍
⚫
⚫
❍
⚫
◐
❍
⚫
❍
Al Masoud et al. [ 110 ]
SoK
❍
❍
❍
⚫
❍
❍
❍
◐
◐
◐
❍
❍
⚫
❍
Agent economy (8)
Zhang et al. [ 14 ]
SoK
⚫
◐
◐
◐
⚫
◐
❍
◐
⚫
⚫
◐
◐
◐
❍
Li et al. [ 17 ]
Ana
⚫
◐
◐
❍
⚫
◐
⚫
⚫
⚫
⚫
⚫
◐
⚫
⚫
Appendix
TABLE VIII: Analysis coverage in closely related work.
Obligation
Evidence / contract
Preservation check
Failure boundary
P5
Price 10; cP≤10≤vR ; outside options 0
S5: transfers, cost and delivery assumptions
Extra fees/nonpayment defeat utility bounds
P6
Escrow prefunded with 10
S5: capped payout; no double settlement
Subsidies and extra rewards need accounting
P9a
Sound check of the committed sorting relation
S5: authenticity, validity, all task bindings
No verifier: hole; valid receipt ignored: break
P9b
Justified acceptance path; callable verifier
Access, cost, schedule: acceptance by 12
Silence or terminal refund can block access
P13
Bounded delivery/scheduling; valid receipt
Payment by 16; S6: matching terminal record
Eventual delivery gives no finite deadline
P15
Authenticated identities and consequential acts
S6: signatures and task-specific attribution
Signatures do not establish verdict correctness
Appendix
TABLE IX: Evidence and preservation checks for the reference task.
Model
Correct tool result
Correct usable submission
S
N
Qwen3.5 4B
8/8
0/8
8/8
Gemma 4 E4B
8/8
8/8
8/8
GPT-5.6 Sol
8/8
8/8
8/8
MiniMax M2.5
8/8
8/8
8/8
GLM-5.3
8/8
8/8
8/8
Appendix
TABLE X: Agent submission and settlement outcomes.
In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability problem for multi-agent systems: how should generation, critique, coordination, and human judgment be organized when no component can certify the final result? We address this problem through pAI-Econ-claude, a gated, human-in-the-loop multi-agent architecture for AI-assisted economic theory development. Agents coordinate through a shared workspace of inspectable intermediate records; specialized gates diagnose targeted failure modes and recommend loopbacks without certifying correctness; and human checkpoints retain authority over decisions that are costly to reverse. We evaluate the architecture on five matched economic-theory tasks against an ungated baseline. Two evaluators blinded to configuration agreed on all five pairwise rankings, preferring the gated architecture in four tasks and the baseline in one. Mean failure severity fell from 1.58 to 1.16, while overall usefulness rose from 2.60 to 3.10. The largest observed gain occurred when a reality check rejected a false market-structure premise and a proof review prompted revision of a false welfare claim. The negative case shows that scaffolding can also compress an economically important mechanism too aggressively. The results support a bounded claim: gated oversight improves the auditability of AI-assisted economic theory without substituting for formal verification, and the allocation of irreversible human judgment is a more informative design variable than pure agent autonomy. The workflow is publicly available at https://github.com/maxwell2732/pAI-Econ-claude.
Chen Zhu, Xiaolu Wang, Weilong Zhang
China Agricultural University · University of Cambridge
Agent-runtime systems emit traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records, but those containers do not necessarily answer governance questions about a specific decision. DEMM-Bench is a cross-regime benchmark for agent-runtime governance-evidence sufficiency, grounded in the Decision Evidence Maturity Model (DEMM): it measures whether records across eight evidence regimes are sufficient to reconstruct decision-level properties rather than merely present. The benchmark normalizes the regimes through adapters, asks property questions over actor, authority, action, policy, decision basis, resource touch, lifecycle context, and verification strength, and applies eight deterministic degradation conditions. Across 64 manuscript cases, trace-present and schema-present baselines overclaim on 75% of cases, ledger-present overclaims on 50%, and the redacted property-level candidate scorer has zero overclaim with 56.25% mean Property Sufficiency Accuracy. The deposited package provides the 64-case dataset, construction-oracle labels, baselines, and adapters, supporting reproducible evaluation of decision-evidence maturity across heterogeneous agent-runtime evidence substrates.
Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies completion evidence, and selectively invalidates affected work. Across controlled comparisons, governed and direct workflows often reached the same outcomes, but only the governed path consistently preserved governing evidence, refused unsupported closure, and limited recovery to dependent tasks. A role-separated transfer challenge then failed: a deterministically enforced completeness contract severely over-blocked synthetic packets produced outside its authoring context. These results do not establish Matrix as a general accuracy enhancer; they support its primary role as an institutional integrity layer for making agentic work auditable and independently verifiable.
Jesus Salas
Independent Researcher · The author is employed by Microsoft.