The Last Human Gate: Forward Deployed Engineering for Governance Automation
Abstract
Enterprise governance requires decisions, evidence, and accountable authority; it does not require every review task to retain its current human implementation. We develop a task-substitution framework for Digital Governance Frameworks (DGF), treating each gate as an executable contract. Substitution requires sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted. We derive a residual-work threshold and show why automating most cases can still increase labor. Forward deployed engineering connects these conditions to an architecture for agents, rule engines, evidence services, and escalation. DGF-Bench supplies controlled evidence from 300 synthetic projects and 899 evaluable model-project runs. Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%; complete-route success is 76.92%, 42.33%, and 24.67%. A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the comparison in execution of a supplied decision kernel. Evidence audits and 135 repeated runs distinguish correct decisions from reliable execution. A document counterexample establishes an information-sufficiency obstruction. These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software. The framework specifies a workforce test based on the complete human effort required at fixed output and quality; the present measurements concern review performance. Sources, dossiers, traces, and analyses are public.
Figures & tables
| Claim | Observable test | Evidence in this paper |
|---|---|---|
| Contract execution | Decision, findings, actions, evidence, and mandate meet fixed criteria | Original runs and deterministic control |
| Workflow reliability | Every required gate succeeds; repeated trajectories remain successful | Complete-route scores and 135 additional runs |
| Net task substitution | Accepted output requires fewer total human hours | Accounting conditions and illustrative calculations |
| Ecosystem workforce reduction | Comparable governed output and quality with a lower complete FTE account | Dated hypothesis and a measurement protocol |
| Benchmark | Central task | Relation to this study |
|---|---|---|
| WorkArena | Execute knowledge-work tasks through enterprise web interfaces | Interface operation can support evidence collection, but is not the gate-review endpoint here |
| ITBench | Address operational IT automation scenarios | Operational execution and remediation differ from reviewing whether a project meets a governance contract |
| -bench | Use tools while following domain policies in a user interaction | Motivates repeated reliability and policy compliance; DGF-Bench follows specialist reviews through a dossier route |
| DGF-Bench | Produce a disposition with findings, actions, observed evidence, and authorization | Measures complete review records and whole-route success under explicit supplied policies and facts |
| Gate | Review question | Examples of material examined |
|---|---|---|
| Procurement | Does the selected supplier satisfy purchasing requirements? | Offers, mandatory criteria, due diligence, sanctions, and three-year cost |
| Legal | Are the contractual prerequisites and signing authority present? | Data-processing agreement, liability, exit assistance, and signing authority |
| Compliance | Does the proposal satisfy the applicable compliance checks? | Required impact assessment, residency, regulatory mapping, and audit trail |
| Security | Does the design provide the protections required for its exposure and data? | Identity controls, private endpoints, monitoring, and vulnerability findings |
| IT | Does the proposal fit the technology and operating environment? | Catalog status, existing capability, capacity, support ownership, and change records |
| Architecture | Is the proposed system consistent with the declared design constraints? | Network overlap, interfaces, data ownership, latency, and reversibility |
| Property | Potential advantage | Condition that can defeat it |
|---|---|---|
| Recurring dossiers and outputs | Reusable extraction and decision interfaces | Material facts remain tacit or inaccessible |
| Explicit policies and mandates | Executable checks and bounded delegation | Conflicting rules or discretionary commitments |
| Recorded gate handoffs | Shared evidence and traceable dependencies | Work is repeated or lost between teams |
| Repeated review populations | Integration cost spread across cases | Low volume or frequent policy change |
| Observable acceptance criteria | Regression tests and monitored service | Evaluation rewards format instead of valid work |
| Quantity | Required observation | Interpretation |
|---|---|---|
| Review visits and baseline hours | Scale and baseline intensity | |
| Routed cases and their baseline hours | Case coverage versus retained labor | |
| Complete exception effort | Assistance or inflation on the residual | |
| Ordinary review and audit effort | Human work remaining on the standard path | |
| Additional correction and rework | Work created beyond the recorded paths | |
| Allocated recurring support | Human upkeep in both regimes |
| Scenario | Visits | Support h/visit | ||||
|---|---|---|---|---|---|---|
| Integrate | 5 | 0.60 | 1.4 | 1.0 | 13.20 | 1.451 |
| Build | 95 | 0.14 | 2.0 | 0.8 | 0.69 | 0.360 |
| Integrate at higher volume | 95 | 0.60 | 1.4 | 1.0 | 0.69 | 0.930 |
| Deliverable | Acceptance criterion | Human-work location |
|---|---|---|
| Source interface | Recover required premises or report their absence; preserve source identity | Preparation and verification; recurring connector upkeep in |
| Policy and mandate service | Apply the specified version; reject unauthorized commitments; retain unresolved conditions | Ordinary checks in ; exceptional decisions in |
| Evidence and handoff record | Findings retain observed support and their open obligations through downstream reviews | Repeated review in or ; additional repair in |
| Operating ledger | Record case work and shared support across teams and suppliers | Case work in ; shared recurring effort in |
| Outcome | Gemini | Luna | DeepSeek | Rules |
|---|---|---|---|---|
| Evaluable projects | 299 | 300 | 300 | 300 |
| Evaluable gates | 1,694 | 1,700 | 1,700 | 1,700 |
| Correct dispositions | 1,694 | 1,668 | 1,619 | 1,700 |
| Strict gate successes | 1,609 | 1,416 | 1,261 | 1,700 |
| Strict gate success (%) | 94.98 | 83.29 | 74.18 | 100.00 |
| Complete routes | 230 | 127 | 74 | 300 |
| Gate | DeepSeek | Gemini | Luna |
|---|---|---|---|
| Architecture | 75.00 | 85.43 | 89.50 |
| Compliance | 88.00 | 100.00 | 96.00 |
| General | 68.67 | 87.96 | 61.67 |
| IT | 87.00 | 100.00 | 97.67 |
| Legal | 68.50 | 99.50 | 80.50 |
| Procurement | 43.00 | 96.00 | 33.00 |
| Recorded quantity | Gemini | Luna | DeepSeek |
|---|---|---|---|
| Approval calls | 864 | 394 | 216 |
| Rejected calls | 466 | 40 | 14 |
| Used conditional decisions | 398 | 353 | 202 |
| Eligible gates, all scored runs | 398 | 398 | 397 |
| Used on matched opportunities / 391 | 391 | 347 | 201 |
| Condition | Policy representation | Permitted evidence |
|---|---|---|
| A | Executable definitions | Authoritative structured snapshots |
| B | Semantically equivalent prose | The same snapshots |
| C | The same executable definitions | Documents and operational records sufficient to recover the same facts |
| D | The same prose as B | The same documents and records as C |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Claim or operation | Inspectable record |
|---|---|
| Original metrics and intervals | research/2026-09-dgf-bench/paper_results.json |
| Matched model comparisons | research/2026-09-dgf-bench/paired_comparisons.json |
| Falcon and Meridian traces | research/2026-09-dgf-bench/worked_examples.json |
| Structural evidence audit | research/2026-09-followup/ALL_MODELS_AUDIT.md |
| Repeated trajectories | research/2026-09-followup/repetition_analysis.json |
| Information counterexample | research/2026-09-followup/document_ablation_preflight.json |