Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
Organizations: SynWe Group s.r.o.
Abstract
Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose verdicts are themselves probabilistic; we ask whether a deterministic layer can instead stop contract-detectable handoff defects. We present Maat, a runtime governance layer that validates agent-to-agent handoffs against a versioned workflow contract, or anchor, with no language model in the validation or scoring path. We evaluate it in six controlled domain workflows (6-15 agents, 522 trials) with injected data-level defects and a deterministic seven-check rubric. Version 1 reported gains in all six workflows (2.9-26.5%). A post-publication audit found that three benchmark scorers credited any early halt as a prevented defect. On paired trials where the governed run completed or halted on a finding attributable to a verified defect, the rubric score changes by +7.7% to +29.1% in five workflows and is flat in software development; model-call cost falls 17-53% where attributable halts occur early. A hand review of all 94 governed-arm halts found 35 false alarms (37%), caused by validator defects rather than model behaviour; counting those halts as failed work, the governed arm scores below the ungoverned arm in four of six workflows. The results support deterministic handoff validation for contract-expressible defects and show that validator configuration and halt attribution must themselves be tested; they do not establish universal correctness, hallucination detection, or model-independent effectiveness.
Figures & tables
| Benchmark | Agents | Trials | Workflow |
|---|---|---|---|
| B2B financial | 6 | 75 | lead-to-invoice |
| Hospital triage | 10 | 75 | intake-to-disposition |
| Software development | 12 | 75 | specification-to-release |
| Enterprise discovery | 12 | 90 | process discovery with amendments |
| E-commerce agency | 15 | 135 | multi-marketplace, two concurrent clients |
| Health insurance | 10 | 72 | EU claims adjudication |
| correctness | Cost | ||||||
|---|---|---|---|---|---|---|---|
| Benchmark | v1 | Paired | Lower bound | Halts (unattr.) | v1 | Paired | |
| B2B financial | +26.5% | +29.1% | 17/25 | -15.9% | 20 (8) | -50% | -47% |
| Hospital triage | +12.4% | +11.1% | 9/25 | -60.8% | 22 (16) | -49% | -21% |
| Software development | +3.8% | +0.8% | 19/25 | -23.1% | 7 (6) | -8% | -1% |
| Enterprise discovery | +7.7% | +7.7% | 30/30 | +7.7% | 5 (0) | 0% | 0% |
| E-commerce agency | +10.7% | +10.6% | 44/45 | +8.3% | 18 (1) | -17% | -17% |
| Case | Observable violation | Why the boundary is checkable |
|---|---|---|
| Marketplace price drift | Same keyed SKU carries conflicting prices | Stable entity, divergent tracked value |
| Payout inflation | Published payout differs from contract-derived expected value | Numeric decision conflicts with encoded authority |
| VAT/tax mismatch | Transaction classification conflicts with jurisdiction rule | Jurisdiction and transaction class are explicit state |
| Claimant identity drift | Policyholder identifier changes across handoffs | Identity should remain invariant through adjudication |
| Data residency | Processing destination violates the encoded rule | Processing location is explicit policy-bounded state |
| Revenue mis-report | Aggregate does not reconcile to source transactions | Deterministic reconciliation exposes the mismatch |
| Stronger fit | Weaker fit |
|---|---|
| Stable policy or contract exists | No authoritative state can be encoded |
| Identity/value must survive many handoffs | Agents already independently reconstruct the same facts |
| Violation has financial, legal, safety, or audit impact | Errors are harmless or cheaply reversible |
| Downstream model calls are expensive | Workflow is short and inexpensive |
| Auditability and reproducibility matter | Open-ended semantic quality is the main objective |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| MAST failure mode | Maat mechanism |
|---|---|
| Disobey task specification | G2 plus requirement primitives compare output to the anchor |
| Disobey role specification | G4 checks declared role boundaries (public empirical coverage remains limited) |
| Information withholding | G2 required-field completeness |
| Ignored other agent’s input | G7 cross-handoff entity consistency |
| Premature termination / unhealthy execution | G5–G6 survival monitoring and circuit breaking |
| No or incomplete verification | deterministic validation at handoff boundaries |
| Benchmark | Scenario | Mechanism | Producer consumer | Handoff |
|---|---|---|---|---|
| Insurance | Payout inflation | REQ_VALUE_MISMATCH | decision payment | decision/payment |
| Insurance | Claimant identity drift | G7 entity drift | coverage medical | coverage/medical |
| Insurance | Out-of-network provider | REQ_PROVIDER_INELIGIBLE | provider compliance | provider/compliance |
| Insurance | VAT mismatch | REQ_TAX_MISMATCH | compliance legal | compliance/legal |
| Insurance | Data residency | REQ_DATA_RESIDENCY | medical fraud | medical/fraud |
| E-commerce | Discount fabrication | REQ_DISCOUNT_EXCEEDED | marketing inventory | marketing/inventory |