Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Organizations: TextQL
Abstract
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
Figures & tables
| Benchmark | Tables per DB | Rows per DB | Data | Coherent system | Python and ML | Actions graded | Ground truth |
| BIRD ( Li et al., 2023 ) | 7.3 | 549K | Public | ✗ | ✗ | ✗ | Gold SQL |
| Spider 2.0 ( Lei et al., 2025 ) | 52.6 † | n/r | Public | ✗ | ✗ | ✗ | Gold SQL |
| BEAVER ( Chen et al., 2024 ) | 101.5 | n/r | Private | ✓ | ✗ | ✗ | Logged SQL |
| DSBench ( Jing et al., 2025 ) | n/r | n/r | Public | ✗ | ✓ | ✗ | Answer keys |
| DABstep ( Egg et al., 2025 ) | n/r | 138K | Real | ✗ | ✓ | ✗ | Answer keys |
| -bench ( Yao et al., 2025 ) | 3 | 2.8K | LLM-made | ✗ | ✗ | ✓ | Goal state |
| Overall | Score by domain | ||||||||
| Model | Solved (%) | Score | Fcst. | Fraud | Fin. | Dash. | Comp. | Steps | Cost ($) |
| Proprietary | |||||||||
| GPT-6 Astra | 27.6 | 51.8 | 39.5 | 48.0 | 85.8 | 71.0 | 85.7 | 23 | 2.71 |
| GPT-6.1 Sol | 24.8 | 49.5 | 38.4 | 47.6 | 87.1 | 63.4 | 82.8 | 24 | 0.49 |
| GPT-6 Sol | 17.6 | 36.8 | 27.9 | 35.4 | 75.8 | 37.3 | 75.7 | 38 | 0.85 |
| GPT-6 Luna | 7.6 | 15.0 | 2.1 | 19.1 | 56.1 | 10.1 | 42.9 | 39 | 0.06 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Donor | Used for | Role |
| Geography | ||
| Overture Places | The 45,834 restaurants, with names and locations | I |
| DOHMH inspections | Cuisine, license, grades, openings and closures since 2017 | I |
| Overture Addresses | 1.07M residential address points | I |
| MapPLUTO | Building use, units, and floor area of 857k tax lots | I |
| TLC taxi zones | The 260 zones that partition the city | I |
| Category | Task | Examples | Cutoff |
| Before the shock | Forecast April’s true-ups and hours from the first quarter, before any true-up has been paid | fc-08-true-ups-march , fc-13-april | March |
| Just after | Recover the true-up rule from the first weeks under it and forecast pay, hours, true-ups or margin for May to the third quarter | fc-08 , fc-10 , fc-11 , fc-12 (Table L ) | April or May |
| Later in the year | Forecast July’s true-ups, third-quarter courier costs, and December’s no-courier cancellations through further dispatcher and pay-card changes | fc-10t2 , fc-13-july , fc-19 , fc-16 | June to November |
| Regulatory checks | Find the pay periods that paid couriers short, restate pay per connected hour on the rule’s basis, and size back pay after a complaint | pvl-01 , vd-e3 , fc-14-back-pay | August to December |
| Dashboards | Rebuild margin and courier-cost tiles whose true-up component starts in April | ds-22 to ds-25 , ds-33 , ds-36t , erp-01 | December |
| Orders & parties 51 tables · 723 M rows | |||
| Table | Cols | Rows | |
| Order Management 6 tables · 580 M rows | |||
| OE_ORDER_HEADERS_ALL | 41 | 81 M | |
| OE_ORDER_LINES_ALL | 24 | 446 M | |
| OE_ORDER_SOURCES | 12 | 40 | |
| OE_PRICE_ADJUSTMENTS | 35 | 52.9 M | |
| Finance (cont.) 79 tables · 4.18 B rows | |||
| Table | Cols | Rows | |
| Subledger Accounting 7 tables · 1.33 B rows | |||
| XLA_AE_HEADERS | 22 | 167 M | |
| XLA_AE_LINES | 25 | 413 M | |
| XLA_DISTRIBUTION_LINKS | 7 | 413 M | |
| XLA_EVENTS | 18 | 167 M | |
| Mode | Scored by | Tasks |
| forecast | Weighted interval score, on a scale set by a reference forecast | 72 |
| cost_set | Cost saved by bans or holds (Equation 1 ) | 43 |
| data_source | Structural contract, then every value of the tile | 36 |
| presence | Whether a filing of that kind was made, such as a written note | 34 |
| id_set | Overlap of a set of IDs with the key (F1) | 28 |
| keyed_value | Values per key, each within a tolerance | 22 |
| Task (expectations) | Point | |
| fc-08 March, hinted and not (2) | The minimum-pay rule applied to the March payout periods, scaled to May’s | Rule: spread of the March weeks |
| fc-13 April, July (2) | The aggregate floor gap from the plans in the prompt and last month’s pay per delivery | Rule: last month’s change in pay per delivery |
| fc-10h June hours (1) | The dispatcher’s plan: five periods at the mean of the last four | Hand |
| fc-11h June pay (1) | FP&A’s plan: five periods at the mean of the last eight | Hand |
| fc-10h , fc-11h Q3 (2) | The latest week carried through thirteen periods, less 1.3% a week for hours; chosen after looser plans were beaten | Hand |
| fc-14 back pay (1) | Engineering’s estimate: dropped trip hours priced at $19.56 in the weeks the top-up ran | Hand |
| Source of the database | Questions | Share |
| BigQuery public data, Google | 341 | 62.3% |
| BigQuery public data, third party | 39 | 7.1% |
| Local SQLite files (Kaggle, textbook and vendor samples) | 135 | 24.7% |
| Snowflake Marketplace | 18 | 3.3% |
| Other public dumps (Meta Kaggle, WideWorldImporters, CDC) | 14 | 2.6% |
| Model | Score | Solved (%) |
| Claude Opus 5.5 | 59.5 [52.0, 67.1] | 34.8 [27.3, 43.3] |
| GPT-6 Astra | 51.8 [43.8, 59.8] | 27.6 [20.8, 35.3] |
| Claude Sonnet 5.5 | 51.8 [44.5, 59.0] | 28.6 [21.5, 36.5] |
| GPT-6.1 Sol | 49.5 [41.8, 57.0] | 24.8 [18.4, 32.0] |
| GPT-6 Sol | 36.8 [29.5, 44.0] | 17.6 [12.3, 23.8] |
| Kimi K3 | 28.4 [21.0, 36.7] | 14.3 [9.5, 19.7] |
| Model | Tasks | Low | Medium | High | Extra-high |
| Claude Opus 5.5 | 210 | 32.6 | 49.6 | 55.5 | 59.5 |
| GPT-6 Astra | 210 | 33.3 | 45.1 | 47.6 | 51.8 |
| Claude Sonnet 5.5 | 210 | 19.5 | 25.4 | 40.3 | 51.8 |
| GPT-6.1 Sol | 210 | 24.6 | 35.9 | 43.6 | 49.5 |
| GPT-6 Sol | 210 | 19.5 | 28.2 | 30.7 | 36.8 |
| Kimi K3 | 210 | 12.5 | – | 22.6 | 28.4 |
| Published | Settings | Member windows over the key |
| The right table | 3 | 0 |
| Right rows, membership from contract dates cut at termination | 12 | 7,664 |
| Right rows, membership from contract dates uncut | 12 | 36,824 |
| Right rows, other errors | 6 | |
| Wrong key column | 4 | |
| Nothing | 10 |
| GPT-6 Astra | Claude Opus 5.5 | |
| Score | 0 | 99 |
| Budget met | yes | yes |
| Net saving | $86,281 | $3,089,763 |
| Saving of a uniform cut | $402,037 | $402,037 |
| Storefronts searched | Mean score |
| 1,500 to 15,000 orders | 28.6 |
| 1,500 to 45,000 orders | 16.8 |
| Every storefront | 21.6 |
| Every storefront, with fewer hints about the mechanism | 15.5 |
| Task | Runs | No-change reference | FP&A’s plan |
| June courier base pay | 47 | 84.7 | 6.0 |
| BigQuery $ per task | ||||||
| Model | Mean | Median | TiB per task | Share of cost (%) | Unsolved (%) | $ per solve |
| Claude Opus 5.5 | 0.82 | 0.27 | 0.132 | 15 | 9 | 2.37 |
| GPT-6 Astra | 0.61 | 0.13 | 0.097 | 18 | 15 | 2.20 |
| Claude Sonnet 5.5 | 0.68 | 0.28 | 0.109 | 15 | 30 | 2.39 |
| GPT-6.1 Sol | 0.60 | 0.16 | 0.096 | 55 | 22 | 2.43 |
| GPT-6 Sol | 0.80 | 0.18 | 0.128 | 49 | 42 | 4.53 |