cs.CLOct 1, 2026

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Authors: Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Organizations: TextQL

Abstract

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

    Aug 11, 2026Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3Data Science AgentsMemoryagentbench

  2. AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

    May 22, 2026Darek Kleczek, Fuheng Zhao, Alexander W. Lee +4Enterprise AnalyticsLatent World Models

  3. DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness

    Jul 30, 2026Debin Meng, Jiaming Yang, Zefang Zong +4Data Science AgentsMle-Bench Lite