EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
Organizations: Shandong University · Zhongguancun Academy · Beijing Institute of Technology · Tsinghua University
Abstract
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
Figures & tables
| Domain | Task | Single-agent | Multi-agent | |||||||
| CoT | Self-refine | Reflexion | AMEM | Debate | Discussion | DC | GEPA | ACE | ||
| Information Extraction | All | 0.824 | 0.826 | 0.791 | 0.857 | 0.798 | 0.830 | 0.801 | 0.846 | 0.853 |
| Numerical Calculation | Easy | 0.840 | 0.825 | 0.816 | 0.847 | 0.776 | 0.624 | 0.834 | 0.836 | 0.825 |
| Middle | 0.730 | 0.697 | 0.712 | 0.711 | 0.588 | 0.527 | 0.700 | 0.736 | 0.712 | |
| Hard | 0.500 | 0.500 | 0.500 | 0.563 | 0.375 | 0.375 | 0.438 | 0.563 | 0.500 | |
| Domain Knowledge | Easy | 0.925 | 0.906 | 0.981 | 0.925 | 0.906 | 0.925 | 0.793 | 0.925 | 0.925 |
| Domain | Task | Single-agent | Multi-agent | |||||||
| CoT | Self-refine | Reflexion | AMEM | Debate | Discussion | DC | GEPA | ACE | ||
| Information Extraction | All | 0.863 | 0.852 | 0.837 | 0.871 | 0.861 | 0.870 | 0.822 | 0.863 | 0.858 |
| Numerical Calculation | Easy | 0.860 | 0.851 | 0.850 | 0.931 | 0.851 | 0.864 | 0.922 | 0.943 | 0.963 |
| Middle | 0.742 | 0.691 | 0.736 | 0.727 | 0.718 | 0.731 | 0.863 | 0.842 | 0.851 | |
| Hard | 0.500 | 0.500 | 0.500 | 0.500 | 0.500 | 0.500 | 0.500 | 0.563 | 0.437 | |
| Domain Knowledge | Easy | 0.887 | 0.887 | 0.944 | 0.962 | 0.906 | 0.906 | 0.813 | 0.921 | 0.925 |
| Backbone | Dimension | Single-agent | Multi-agent | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CoT | Self-refine | Reflexion | AMEM | Debate | Discussion | DC | GEPA | ACE | ||
| DeepSeek-V3 | Structure | 7.26 | 6.54 | 6.90 | 7.40 | 6.60 | 6.99 | 8.15 | 8.08 | 7.68 |
| Quant. | 6.58 | 5.58 | 5.92 | 6.99 | 5.67 | 6.06 | 7.90 | 8.08 | 7.30 | |
| Business | 7.30 | 6.56 | 6.93 | 7.44 | 6.73 | 7.01 | 8.15 | 8.40 | 7.70 | |
| Comm. | 7.60 | 6.97 | 7.33 | 7.92 | 6.94 | 7.27 | 8.00 | 8.59 | 8.07 | |
| Overall | 7.12 | 6.38 | 6.75 | 7.37 | 6.45 | 6.80 | 8.00 | 8.28 | 7.63 | |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Train | Valid | Test | Total |
|---|---|---|---|---|
| Structured Reasoning | ||||
| CodeFinQA | 4,409 | 200 | 788 | 5,397 |
| CodeTAT-QA | 2,654 | 200 | 288 | 3,142 |
| ConvFinQA | 133 | - | 132 | 265 |
| FinCode | 7 | 2 | 47 | 56 |
| finer | 1,000 | 500 | 441 | 1,941 |
| Domain | Task | Single-agent | Multi-agent | |||||
| CoT | Self-refine | Reflexion | AMEM | Debate | Discussion | GEPA | ||
| Information Extraction | All | 0.874 | 0.876 | 0.856 | 0.864 | 0.871 | 0.895 | 0.939 |
| Numerical Calculation | Easy | 0.872 | 0.865 | 0.856 | 0.875 | 0.860 | 0.865 | 0.882 |
| Middle | 0.742 | 0.776 | 0.755 | 0.776 | 0.761 | 0.776 | 0.784 | |
| Hard | 0.500 | 0.625 | 0.438 | 0.563 | 0.500 | 0.625 | 0.571 | |
| Domain Knowledge | Easy | 0.962 | 0.962 | 0.981 | 0.943 | 0.962 | 0.962 | 0.962 |
| Domain | Task | Single-agent | Multi-agent | |||||
| CoT | Self-refine | Reflexion | AMEM | Debate | Discussion | GEPA | ||
| Information Extraction | All | 0.903 | 0.898 | 0.875 | 0.907 | 0.897 | 0.911 | 0.932 |
| Numerical Calculation | Easy | 0.826 | 0.821 | 0.701 | 0.842 | 0.841 | 0.845 | 0.843 |
| Middle | 0.730 | 0.746 | 0.642 | 0.755 | 0.755 | 0.746 | 0.752 | |
| Hard | 0.500 | 0.563 | 0.438 | 0.563 | 0.563 | 0.563 | 0.563 | |
| Domain Knowledge | Easy | 0.925 | 0.962 | 0.925 | 0.943 | 0.925 | 0.906 | 0.943 |
| Backbone | Dimension | Single-agent | Multi-agent | |||||
|---|---|---|---|---|---|---|---|---|
| CoT | Self-Refine | Reflexion | AMEM | Debate | Discussion | GEPA | ||
| DeepSeek-V4-Pro | Structure | 7.21 | 6.54 | 6.88 | 7.59 | 6.48 | 7.00 | 7.92 |
| Quant. | 6.53 | 5.50 | 5.82 | 7.20 | 5.51 | 6.03 | 7.68 | |
| Business | 7.22 | 6.54 | 6.78 | 7.68 | 6.55 | 6.92 | 8.21 | |
| Comm. | 7.51 | 6.95 | 7.21 | 8.15 | 6.78 | 7.19 | 8.51 | |
| Overall | 7.06 | 6.33 | 6.63 | 7.58 | 6.29 | 6.74 | 8.10 | |
| Dimension | Pearson | Spearman | MAE | Within-1 | Human ICC |
|---|---|---|---|---|---|
| Structure | 0.873 | 0.743 | 0.752 | 90.4% | 0.861 |
| Quantitative Reasoning | 0.840 | 0.797 | 0.929 | 67.4% | 0.782 |
| Business Sense | 0.864 | 0.745 | 0.712 | 84.4% | 0.834 |
| Communication | 0.899 | 0.656 | 0.900 | 72.0% | 0.876 |
| Overall | 0.887 | 0.736 | 0.679 | 85.0% | 0.892 |
| Method | Acc. | Calls | Calls/ sample | Total tokens | Tokens/ sample | Memory/history access | Prompt evolution |
|---|---|---|---|---|---|---|---|
| CoT | 0.98 | 50 | 1.00 | 54,292 | 1,086 | None | None |
| AMEM | 0.98 | 50 | 1.00 | 52,502 | 1,050 | Online memory from previous samples only | None |
| Self-Refine | 0.92 | 103 | 2.06 | 130,014 | 2,600 | Current-sample trajectory only | None |
| Reflexion | 0.98 | 102 | 2.04 | 126,558 | 2,531 | Current-sample trajectory only | None |
| Debate | 0.92 | 150 | 3.00 | 206,305 | 4,126 | Current-sample discussion only | None |
| Discussion | 0.94 | 200 | 4.00 | 228,917 | 4,578 | Current-sample discussion only | None |