Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation
Figures & tables
Figure 1: Two gaps between public benchmarks and enterprise decisions. Capability benchmarks measure what a model can do on portable tasks. Enterprises must also establish transfer to their own workflows and ground truth (Gap 1) and convert quality metrics into a value-and-risk case a governance body can act on (Gap 2).
Figure 2: eVal architecture. Four evaluation layers feed a governance decision; the governance spine consumes artefacts from every layer.
Figure 3: Evaluation-intensity matrix (Eq. 3 ). Markers place the pilot use cases: B credit memos (A1,C3), A control assessment (A2,C3), C procedure transformation (A1,C2).
Family
Metric
Dir.
Grading
Level
Fidelity
Hallucination rate (Eq. 4 )
↓
H,J
A1
Citation precision (lenient/strict)
↑
H,J
A1
Citation recall
↑
J
A1
Utility
Key-element capture
↑
H,J
A1
Irrelevance rate
↓
H,J
A1
Completeness (scope attempted)
↑
A
A2
Table 2: Metric catalogue. Grading: H human, J calibrated judge, A automatic. Level: lowest autonomy level at which the metric is gated.
Figure 4: Two-tier gating. Decisions use the conservative CI bound; an interval straddling both thresholds calls for more samples, not a verdict.
A: Control assessment
B: Credit-memo drafting
C: Procedure transformation
Level / tier / E
A2 / C3 / E3
A1 / C3 / E2
A1 batch / C2 / E2
Inputs
Process–risk–control matrices; control standards
Filings, broker research, internal spreads
Existing procedures (30k corpus, 40k users)
Output
Per-control compliance class + recommendation
One of 13 memo sections, cited
Standardised rewrite
Models
Gemini 2.5 Flash / Pro
Gemini 2.5, Llama 4
Pipeline v1, v2
Gated metrics
KEC, IR, HR, guidance P/R, pass^4
CP, KEC, IR, HR, guidance P/R
Δeff
Status
Charter set; grading in progress
Human-graded results
Estimated-baseline results
Table 3: Pilot portfolio.
Metric
θout
θin
Gemini 2.5
Llama 4
Gate
Citation precision (lenient)
> 65
> 70
88.0
76.0
both Scale
Key-element capture
> 50
> 55
99.0
96.0
both Scale
Overall accuracy (composite)
–
–
93.5
86.0
not gated
Irrelevance rate ↓
< 50
< 55 †
23.0
28.0
both Scale
Hallucination rate ↓
< 10
< 5
1.6
3.2
both Scale
Guidance precision
> 75
> 80
98.0
93.0
both Scale
Table 4: Use Case B human-graded results. ↓ lower is better. Gate column applies Eq. 9 to point estimates for illustration; CIs to be added.
Figure 5: Use Case B against two-tier thresholds. Dashed/dotted lines mark outer/inner thresholds.
Figure 6: Threshold-normalised scores νj (Use Case B). Origin =θout , unit circle =θin , cap 3. Equal-weight and fidelity-weighted EI rank the systems differently.
Figure 7: Use Case C estimated effort by document and pipeline version.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Evaluation lifecycle with re-evaluation feedback loop.
Figure 9: Autonomy levels A0–A4 and the metric families they require.
Dimension
Assistive (A0–A1)
Agentic (A3–A4)
Evaluation implication
Task control
User-led
System plans and decomposes
Plan validity, stopping rule, recovery
Context
Prompted / uploaded
Dynamic retrieval and tool calls
Evidence coverage, retrieval and tool-call correctness
Generative AI systems achieve impressive performance on standard benchmarks yet fail to deliver real-world utility, a disconnect we identify across 28 deployment cases spanning education, healthcare, software engineering, and law. We argue that this benchmark utility gap arises from three recurring failures in evaluation practice: proxy displacement, temporal collapse, and distributional concealment. Motivated by these observations, we argue that generative AI evaluation requires a paradigm shift from static benchmark-centered transparency toward stakeholder, goal, and context-conditioned utility transparency grounded in human outcome trajectories. Existing evaluations primarily characterize properties of model outputs, while deployment success depends on whether interaction with AI improves stakeholders' ability to achieve their goals over time. The missing construct is therefore utility: the change in a stakeholder's capability induced through sustained interaction with an AI system within a deployment context. To operationalize this perspective, we propose SCU-GenEval, a four-stage evaluation framework consisting of stakeholder-goal mapping, construct-indicator specification, mechanism modeling, and longitudinal utility measurement. To make these stages practically deployable, we introduce three supporting instruments: structured deployment protocols, context-conditioned user simulators, and persona- and goal-conditioned proxy metrics. We conclude with domain-specific calls to action, arguing that progress in generative AI must be evaluated through measurable improvements in human outcomes rather than benchmark performance alone.
Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions. Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics. Because the sessions contain internal enterprise content, we do not release the benchmark data; instead, our reusable contribution is the construction and evaluation protocol. On EnterpriseClawBench, the best configuration reaches only 0.663 (Codex with GPT-5.5). These results show that enterprise agent evaluation must report harness--model combinations, artifact delivery, visual quality, cost, runtime, and skill-transfer behavior, rather than collapsing performance into a single score. Code: https://github.com/FrontisAI/EnterpriseClawBench
LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. The benchmark is built around a complete fictional company. It includes product simulators, company-specific internal databases, benchmark questions, and computed answer keys. Industry, company size, business model, application portfolio, and a seed define each company. One seeded entity graph supplies shared company data to simulators of Salesforce, Zendesk, Slack, Gong, and other products. A questionconditioned generator creates the schemas and records for internal databases. It takes shared entities, keys, and values from the same graph before generating database-specific facts. Both mechanisms therefore describe one consistent enterprise estate. Every expected answer is computed from the final records, so grading is exact. Design and answer-key checks validate the internal databases. A realism scorecard and adversarial detector validate the entity graph. Across 23 generated companies, the mean realism score rose from 61.8 to 97.0, with zero records flagged as synthetic. In the reported simulator-track comparison, nine models answered the same 33 questions three times each. Accuracy estimates ranged from 42.4% to 76.8%, and three of 36 pairwise differences remained supported after correction.
Benjamin Gruenbaum, Doron Porat, Assaf Natanzon +3