LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
Figures & tables
Figure 1: Overview of EnterpriseBench. (a) EnterpriseBench organizes enterprise-oriented evaluation into a layered capability structure, progressing from information extraction, domain knowledge, and numerical calculation to complex reasoning. (b)–(d) The benchmark further introduces three interactive decision-making tasks: Consulting for hidden-information business problem solving, the Beer Game for delayed-feedback supply-chain control, and Enterprise Digital Twin for project-based enterprise operation.
Domain
Task
Single-agent
Multi-agent
CoT
Self-refine
Reflexion
AMEM
Debate
Discussion
DC
GEPA
ACE
Information Extraction
All ↑
0.824
0.826
0.791
0.857
0.798
0.830
0.801
0.846
0.853
Numerical Calculation
Easy ↑
0.840
0.825
0.816
0.847
0.776
0.624
0.834
0.836
0.825
Middle ↑
0.730
0.697
0.712
0.711
0.588
0.527
0.700
0.736
0.712
Hard ↑
0.500
0.500
0.500
0.563
0.375
0.375
0.438
0.563
0.500
Domain Knowledge
Easy ↑
0.925
0.906
0.981
0.925
0.906
0.925
0.793
0.925
0.925
Table 1: Performance comparison of single-agent and multi-agent methods across different domains and tasks using DeepSeek-V3 as the backbone model. BeerGame results are reported as total costs in units of ×104 (lower is better), while EDT results are reported in units of 106 (higher is better).
Domain
Task
Single-agent
Multi-agent
CoT
Self-refine
Reflexion
AMEM
Debate
Discussion
DC
GEPA
ACE
Information Extraction
All ↑
0.863
0.852
0.837
0.871
0.861
0.870
0.822
0.863
0.858
Numerical Calculation
Easy ↑
0.860
0.851
0.850
0.931
0.851
0.864
0.922
0.943
0.963
Middle ↑
0.742
0.691
0.736
0.727
0.718
0.731
0.863
0.842
0.851
Hard ↑
0.500
0.500
0.500
0.500
0.500
0.500
0.500
0.563
0.437
Domain Knowledge
Easy ↑
0.887
0.887
0.944
0.962
0.906
0.906
0.813
0.921
0.925
Table 2: Performance comparison of single-agent and multi-agent methods across different domains and tasks using GPT-4.1 as the backbone model. BeerGame results are reported as total costs in units of ×104 (lower is better), while EDT results are reported in units of 106 (higher is better).
Backbone
Dimension
Single-agent
Multi-agent
CoT
Self-refine
Reflexion
AMEM
Debate
Discussion
DC
GEPA
ACE
DeepSeek-V3
Structure
7.26
6.54
6.90
7.40
6.60
6.99
8.15
8.08
7.68
Quant.
6.58
5.58
5.92
6.99
5.67
6.06
7.90
8.08
7.30
Business
7.30
6.56
6.93
7.44
6.73
7.01
8.15
8.40
7.70
Comm.
7.60
6.97
7.33
7.92
6.94
7.27
8.00
8.59
8.07
Overall
7.12
6.38
6.75
7.37
6.45
6.80
8.00
8.28
7.63
Table 3: Performance of different agent methods on the Consulting task across multiple capability dimensions under DeepSeek-V3 and GPT-4.1 backbones.
Figure 2: Comparison of human and LLM scores on 20 Consulting cases, averaging AMEM and CoT outputs per case.
Figure 3: (a) Prompt robustness analysis for Consulting using ACE. Reorder-I and Brief-I denote interviewer prompt variants, while Reorder-J and Brief-J denote judge prompt variants. Scores remain stable across prompt variants, with the overall score changing by at most 0.12 points. (b) Human audit of difficulty annotation. Each point represents one audited task. Difficulty scores assigned by the LLM are strongly aligned with human ratings on the 0–10 scale, with Pearson r=0.96 . (c) Performance comparison of AOA method against baseline approaches on QA tasks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Valid
Test
Total
Structured Reasoning
CodeFinQA
4,409
200
788
5,397
CodeTAT-QA
2,654
200
288
3,142
ConvFinQA
133
-
132
265
FinCode
7
2
47
56
finer
1,000
500
441
1,941
Appendix
Table 4: Statistics of EnterpriseBench tasks across training, validation, and testing splits.
Domain
Task
Single-agent
Multi-agent
CoT
Self-refine
Reflexion
AMEM
Debate
Discussion
GEPA
Information Extraction
All ↑
0.874
0.876
0.856
0.864
0.871
0.895
0.939
Numerical Calculation
Easy ↑
0.872
0.865
0.856
0.875
0.860
0.865
0.882
Middle ↑
0.742
0.776
0.755
0.776
0.761
0.776
0.784
Hard ↑
0.500
0.625
0.438
0.563
0.500
0.625
0.571
Domain Knowledge
Easy ↑
0.962
0.962
0.981
0.943
0.962
0.962
0.962
Appendix
Table 5: Performance comparison of single-agent and multi-agent methods across different domains and tasks using DeepSeek-V4-Pro as the backbone model. BeerGame results are reported as total costs in units of ×104 (lower is better), while EDT results are reported in units of 106 (higher is better).
Domain
Task
Single-agent
Multi-agent
CoT
Self-refine
Reflexion
AMEM
Debate
Discussion
GEPA
Information Extraction
All ↑
0.903
0.898
0.875
0.907
0.897
0.911
0.932
Numerical Calculation
Easy ↑
0.826
0.821
0.701
0.842
0.841
0.845
0.843
Middle ↑
0.730
0.746
0.642
0.755
0.755
0.746
0.752
Hard ↑
0.500
0.563
0.438
0.563
0.563
0.563
0.563
Domain Knowledge
Easy ↑
0.925
0.962
0.925
0.943
0.925
0.906
0.943
Appendix
Table 6: Performance comparison of single-agent and multi-agent methods across different domains and tasks using GLM-5.2 as the backbone model. BeerGame results are reported as total costs in units of ×104 (lower is better), while EDT results are reported in units of 106 (higher is better).
Backbone
Dimension
Single-agent
Multi-agent
CoT
Self-Refine
Reflexion
AMEM
Debate
Discussion
GEPA
DeepSeek-V4-Pro
Structure
7.21
6.54
6.88
7.59
6.48
7.00
7.92
Quant.
6.53
5.50
5.82
7.20
5.51
6.03
7.68
Business
7.22
6.54
6.78
7.68
6.55
6.92
8.21
Comm.
7.51
6.95
7.21
8.15
6.78
7.19
8.51
Overall
7.06
6.33
6.63
7.58
6.29
6.74
8.10
Appendix
Table 7: Dimension-level performance on the Consulting task using DeepSeek-V4-Pro and GLM-5.2 as additional backbone models.
Dimension
Pearson r↑
Spearman ρ↑
MAE ↓
Within-1 ↑
Human ICC (2,3)↑
Structure
0.873
0.743
0.752
90.4%
0.861
Quantitative Reasoning
0.840
0.797
0.929
67.4%
0.782
Business Sense
0.864
0.745
0.712
84.4%
0.834
Communication
0.899
0.656
0.900
72.0%
0.876
Overall
0.887
0.736
0.679
85.0%
0.892
Appendix
Table 8: Transcript-level agreement between the LLM judge and human evaluators on 500 Consulting transcripts. MAE denotes mean absolute error, Within-1 denotes agreement within one point on the 0–10 scale, and ICC (2,3) measures the absolute agreement of the averaged ratings from three human evaluators.
Method
Acc.
Calls
Calls/ sample
Total tokens
Tokens/ sample
Memory/history access
Prompt evolution
CoT
0.98
50
1.00
54,292
1,086
None
None
AMEM
0.98
50
1.00
52,502
1,050
Online memory from previous samples only
None
Self-Refine
0.92
103
2.06
130,014
2,600
Current-sample trajectory only
None
Reflexion
0.98
102
2.04
126,558
2,531
Current-sample trajectory only
None
Debate
0.92
150
3.00
206,305
4,126
Current-sample discussion only
None
Discussion
0.94
200
4.00
228,917
4,578
Current-sample discussion only
None
Appendix
Table 9: Resource and adaptation budgets measured on the same 50 evaluation instances. Calls and tokens include adaptation overhead.
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.
Xinran Zhang, Pengrui Lu, Lyumanshan Ye +1
1Shanghai Innovation Institute · 2Beijing Institute of Technology · 4GAIR Lab +1
Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments.
LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. The benchmark is built around a complete fictional company. It includes product simulators, company-specific internal databases, benchmark questions, and computed answer keys. Industry, company size, business model, application portfolio, and a seed define each company. One seeded entity graph supplies shared company data to simulators of Salesforce, Zendesk, Slack, Gong, and other products. A questionconditioned generator creates the schemas and records for internal databases. It takes shared entities, keys, and values from the same graph before generating database-specific facts. Both mechanisms therefore describe one consistent enterprise estate. Every expected answer is computed from the final records, so grading is exact. Design and answer-key checks validate the internal databases. A realism scorecard and adversarial detector validate the entity graph. Across 23 generated companies, the mean realism score rose from 61.8 to 97.0, with zero records flagged as synthetic. In the reported simulator-track comparison, nine models answered the same 33 questions three times each. Accuracy estimates ranged from 42.4% to 76.8%, and three of 36 pairwise differences remained supported after correction.
Benjamin Gruenbaum, Doron Porat, Assaf Natanzon +3