EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
Organizations: Citigrounp, Inc, London, United Kingdom · Ernst & Young LLP, London, United Kingdom · NVIDIA Corporation, London, United Kingdom
Abstract
Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation
Figures & tables
| Dimension | GDPval | /HAL | CLEAR | KAMI | RAGAS | EV |
|---|---|---|---|---|---|---|
| Firm-specific ground truth | ○ | ○ | ○ | ○ | ◐ | ● |
| Blinded expert grading | ● | ○ | ◐ | ○ | ○ | ● |
| Calibrated LLM judge | ◐ | ○ | ◐ | ○ | ● | ● |
| Reliability (pass^ ) | ○ | ● | ● | ● | ○ | ● |
| Cost / efficiency | ● | ● | ● | ● | ○ | ● |
| Assurance / oversight | ○ | ◐ | ● | ○ | ○ | ● |
| Family | Metric | Dir. | Grading | Level |
|---|---|---|---|---|
| Fidelity | Hallucination rate (Eq. 4 ) | H,J | A1 | |
| Citation precision (lenient/strict) | H,J | A1 | ||
| Citation recall | J | A1 | ||
| Utility | Key-element capture | H,J | A1 | |
| Irrelevance rate | H,J | A1 | ||
| Completeness (scope attempted) | A | A2 |
| A: Control assessment | B: Credit-memo drafting | C: Procedure transformation | |
|---|---|---|---|
| Level / tier / | A2 / C3 / E3 | A1 / C3 / E2 | A1 batch / C2 / E2 |
| Inputs | Process–risk–control matrices; control standards | Filings, broker research, internal spreads | Existing procedures (30k corpus, 40k users) |
| Output | Per-control compliance class + recommendation | One of 13 memo sections, cited | Standardised rewrite |
| Models | Gemini 2.5 Flash / Pro | Gemini 2.5, Llama 4 | Pipeline v1, v2 |
| Gated metrics | KEC, IR, HR, guidance P/R, pass^4 | CP, KEC, IR, HR, guidance P/R | |
| Status | Charter set; grading in progress | Human-graded results | Estimated-baseline results |
| Metric | Gemini 2.5 | Llama 4 | Gate | ||
|---|---|---|---|---|---|
| Citation precision (lenient) | 65 | 70 | 88.0 | 76.0 | both Scale |
| Key-element capture | 50 | 55 | 99.0 | 96.0 | both Scale |
| Overall accuracy (composite) | – | – | 93.5 | 86.0 | not gated |
| Irrelevance rate | 50 | 55 † | 23.0 | 28.0 | both Scale |
| Hallucination rate | 10 | 5 | 1.6 | 3.2 | both Scale |
| Guidance precision | 75 | 80 | 98.0 | 93.0 | both Scale |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Assistive (A0–A1) | Agentic (A3–A4) | Evaluation implication |
|---|---|---|---|
| Task control | User-led | System plans and decomposes | Plan validity, stopping rule, recovery |
| Context | Prompted / uploaded | Dynamic retrieval and tool calls | Evidence coverage, retrieval and tool-call correctness |
| Action | Text only | State-changing actions | Authorisation, reversibility, action-safety gates |
| State | Short-lived | Memory across steps | State consistency; memory poisoning tests |
| Failure mode | Wrong content | Wrong content and wrong action | Measure separately |
| Human role | Reviewer / editor | Approver, supervisor | Review load, intervention rate, escalation quality |
| Phase | AI contribution | Metrics |
|---|---|---|
| Identification | Extract issue facts, stakeholders, corrective-action context | Extraction KEC; omission rate; citation precision |
| Evaluation / scoping | Assess reasonableness of remediation plans; completeness of controls | KEC; HR; recommendation validity |
| Remediation design | Classify control design vs. standard and rule library; recommend fixes | Classification P/R vs. frozen gold; pass^ ; completeness |
| Sustainability | Assess whether controls are implemented and operating | Evidence sufficiency; escalation appropriateness |
| Doc | Title (abbrev.) | Manual | v1 | v2 |
| 101889 | Security word rules | 25 | 3 | 2 |
| 102123 | Do-not-release to any caller | 33 | 10 | n/a |
| 129826 | Inactive and dormant accounts | 33 | 10 | n/a |
| 130293 | Verifications and overrides | 25 | 20 | 2 |
| 130724 | Travel suppression | 33 | 10 | 4 |
| 130763 | Vulnerable adult incident report | 33 | 16 | n/a |
| Field | Instruction to grader | Values |
| Assertive? | Does the sentence make a checkable factual claim? Headings, questions and explicit recommendations are not assertive. | Y / N |
| Supported? | For each claim, can you locate support in the permitted sources? Support from your own knowledge does not count. | Fully / Partially / Not |
| Relevant? | Does the sentence contribute to the instruction for this section? | Y / N |
| Citation used? | Is the cited chunk’s content actually used (strict) or at least reflected (lenient)? | Strict / Lenient / Not |
| Key elements | Tick each gold key element present and correct; note incorrect presentations. | checklist |
| Value-add | Correct, relevant content absent from the human baseline. | 0 / 1 / 2 |
| Layer | Recorded | Metrics | Failure example |
|---|---|---|---|
| Input / context | Source identities, completeness, access policy applied | Completeness; policy-violation rate | Required doc missing; restricted source accessed |
| Planning | Sub-tasks, dependencies, stopping rule | Plan validity (judge); step count | Unsafe plan; non-termination |
| Retrieval / memory | Queries, documents, ranks, evidence used; memory reads/writes | Retrieval precision; evidence coverage; poisoning tests | Correct source not retrieved; stale memory |
| Tool use | Tool, arguments, authorisation, result, retries | Tool-call precision; recovery rate; unauthorised-call rate | Wrong entity; action outside permission |
| Synthesis | Claim–evidence linkage | HR; CP; KEC | Unsupported claim carried forward |
| Human oversight | Checkpoint, intervention reason, action, time | Unplanned intervention rate; escalation appropriateness; | Reviewer misses defect; checkpoint bypassed |
| Artefact | SR 11-7 / SS1/23 | NIST AI RMF + GenAI Profile | EU AI Act (high-risk) | ISO/IEC 42001 |
|---|---|---|---|---|
| Charter (Eq. 1), , , thresholds | Model definition, intended use | Map: context, tolerances | Art. 9 risk mgmt | Cl. 6 planning; impact assessment |
| Gold set, sampling frame | Development data documentation | Measure: test-data representativeness | Art. 10 data governance | Annex A data controls |
| Grading records, , judge calibration | Independent validation evidence | Measure: metric validity, human evaluation | Art. 15 accuracy, robustness | Cl. 9 performance evaluation |
| Gate decision, scorecard | Validation outcome, conditions of use | Manage: risk treatment | Art. 14 human oversight | Cl. 9.3 management review |
| Value-and-risk model | Materiality assessment | Govern: benefit–risk | – | Cl. 6.1 objectives |
| Re-evaluation triggers, monitoring | Ongoing monitoring, change control | Manage: post-deployment monitoring | Art. 72 post-market; Art. 12 logging | Cl. 10 improvement |
| Use case ID; business owner; model owner; second-line reviewer |
| Business outcome and value hypothesis |
| Task specification (versions) |
| Autonomy level , consequence tier , intensity |
| Configuration with versions |
| Baseline workflow and measurement mode (estimated / observed) |
| Gated metrics: estimate, CI, , , , grading mode, |