Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.
Figures & tables
Figure 1: SleuthBench construction. A separate validator remeasures the injected pattern before a deterministic function computes the reference answer.
Family
Templates
Representative injections
Representative tasks
Data quality
8
Corrupt a subset of target values and add a binary column marking those rows; replace some category values with a missing-like label; add a group column with one very rare group.
Identify the marker column or the missing-like label; decide whether there is sufficient statistical evidence to make claims about the rare group.
Feature contribution
9
Permute one numerical feature to break its link to the target; add a smooth inverted-U effect of one feature on the target; add an XOR-shaped effect on the target over two numerical features.
Name the noise feature; name the feature and the value at which its effect peaks; name the interacting pair or its conditional direction.
Table 1: The 17 benchmark question templates. Appendix A (Table 4 ) lists every template with its injection and both question wordings.
Framing
Data quality
Feature contribution
Total
Data science
38
32
70
Business
38
32
70
Total
76
64
140
Table 2: Questions per model, by template family and question framing.
Model
Data quality
Feature contribution
Total
Claude Sonnet 5
77.6
18.8
50.7
Claude Opus 4.8
84.2
29.7
59.3
Claude Fable 5
85.5
45.3
67.1
GPT-5.6 Sol
86.8
71.9
80.0
GPT-5.6 Terra
85.5
51.6
70.0
GPT-5.6 Luna
82.9
34.4
60.7
Table 3: Python-only accuracy (%) by model and template family.
Figure 2: Example of an Empirical Layer artifact from the Bike Sharing dataset. The left panel shows the EBM single-feature effect of temp on predicted bike rentals ( cnt ). The right panel is an excerpt from the JSON file storing this artifact.
Figure 3: Accuracy by template family and model. Solid bars show accuracy with Python alone; hatched segments show the accuracy of Python + Empirical Layer. Also see Table 6 (Appendix E ).
Figure 4: Accuracy by template family and question framing, pooled over all six models. Solid bars show accuracy with Python alone; hatched segments show the accuracy of Python + Empirical Layer. Exact values are listed in Tables 9 and 10 (Appendix E ).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Template
What the injection changes
Question wording (data-science and business)
Data quality
Bad-row indicator
Corrupt a subset of target values and add a binary column whose positive entries mark exactly those rows.
Data-science: You can apply exactly one rule based on a single column to exclude problematic rows before analyzing {OUTCOME_COL} . Which column should the rule be based on? Return a single column name. Business: Some records have already been marked for review in a yes-or-no column. Which column contains that review marker? Return a single column name.
Conditional anomaly indicator
Rewrite 5% of outcomes so that they remain within the observed support but become surprising given similar records. Add the true binary flag and same-prevalence placebo flags.
Data-science: Among the binary 0/1 columns, one indicator has 1 exactly on rows whose {OUTCOME_COL} outcome is conditionally inconsistent with otherwise similar records. The injected outcomes remain within the original target support. Which column is the indicator? Return a single column name. Business: Among the 0/1 columns, one indicator’s 1 values mark records whose {OUTCOME_COL} outcome is inconsistent with otherwise similar records, even though the outcome itself remains within the dataset’s observed range. Which column is that indicator? Return a single column name.
Category target outlier
Choose a sufficiently large category and shift its target values up or down until its mean stands out from the other categories.
Data-science: Is there a category in {FEATURE} whose {OUTCOME_COL} values are unusually high or low compared with the other categories? Return only the category label. Business: Is there a category in {FEATURE} whose {OUTCOME_COL} values stand out from the other categories? Return only the category label.
Conditional bad rows
Change the targets of a small set of rows so they remain plausible in the overall target distribution but are unusual given their feature values.
Data-science: Some rows have {OUTCOME_COL} values that are plausible globally but unusual given their feature values. Identify the {N_INJECT_SAMPLES} affected rows and return their row_id values as a comma-separated list. Business: There are {N_INJECT_SAMPLES} rows in this data with an unusual pattern in {OUTCOME_COL} . Return the row_id values as a comma-separated list.
Underpowered group evidence
Add a categorical group column with two categories: a common group A and a very rare group B; leave the target unchanged so the group comparison has a wide interval.
Data-science: For categorical feature {GROUP_COL} , is there enough statistical evidence to distinguish group A from group B with respect to {OUTCOME_COL} ? Return one of: enough evidence, not enough evidence. Business: For the categorical feature {GROUP_COL} , is the evidence strong enough to say the groups behave differently with respect to {OUTCOME_COL} ? Return one of: enough evidence, not enough evidence.
Semantic missing label
Replace a small number of values in a categorical feature with a distinct missing-like label such as “Unknown” or “Not reported.”
Data-science: Which category label in {FEATURE} appears to encode missing, unknown, or not-reported values? Return only the label. Business: Are there any labels in {FEATURE} that we need to pay attention to? If yes, return only the label. If no, return No.
Appendix
Table 4: The 17 benchmark templates and their injections, with the data-science and business question wording from the public template registry. Braced placeholders are filled when each instance is constructed.
Dataset
Source rows
Target
Reference
AI4I 2020 Predictive Maintenance
10,000
Machine failure
( AI4I, 2020 )
Bike Sharing
17,379
Rental count
( Fanaee-T, 2013 )
California Housing
20,640
House value
( scikit-learn, n.d. )
King County Housing
21,613
Sale price
( King County Housing, n.d. )
Metro Interstate Traffic Volume
48,204
Traffic volume
( Hogue, 2019 )
Steel Industry Energy Consumption
35,040
Energy use
( Sathishkumar V E et al., 2021 )
Appendix
Table 5: Source datasets. Each experimental base table has 1,000 rows before injection. Source row counts refer to the full prepared snapshots.
Data quality
Feature contribution
Total
Model
P
P+E
Δ
P
P+E
Δ
P
P+E
Δ
Claude Sonnet 5
77.6
77.6
0.0
18.8
53.1
+34.4
50.7
66.4
+15.7
Claude Opus 4.8
84.2
84.2
0.0
29.7
65.6
+35.9
59.3
75.7
+16.4
Claude Fable 5
85.5
88.2
+2.6
45.3
75.0
+29.7
67.1
82.1
+15.0
GPT-5.6 Sol
86.8
85.5
-1.3
71.9
78.1
+6.3
80.0
82.1
+2.1
GPT-5.6 Terra
85.5
85.5
0.0
51.6
82.8
+31.3
70.0
84.3
+14.3
Appendix
Table 6: Accuracy (%) by model and template family under P and P + E.
Model or family
N
P: C/P/I
P+E: C/P/I
Claude Sonnet 5
140
71/12/57
93/8/39
Claude Opus 4.8
140
83/9/48
106/8/26
Claude Fable 5
140
94/8/38
115/10/15
GPT-5.6 Sol
140
112/9/19
115/9/16
GPT-5.6 Terra
140
98/9/33
118/6/16
GPT-5.6 Luna
140
85/8/47
97/10/33
Appendix
Table 7: Grade distributions by model and by family. C/P/I denotes Correct / Partial / Incorrect .
Model
Framing
N
P (%)
P+E (%)
Δ (pp)
Claude Sonnet 5
Business
70
47.1
67.1
+20.0
Claude Sonnet 5
Data science
70
54.3
65.7
+11.4
Claude Opus 4.8
Business
70
52.9
74.3
+21.4
Claude Opus 4.8
Data science
70
65.7
77.1
+11.4
Claude Fable 5
Business
70
62.9
84.3
+21.4
Claude Fable 5
Data science
70
71.4
80.0
+8.6
Appendix
Table 8: Accuracy by model and question framing. Every model is evaluated on both framings of the same 70 dataset–template combinations.
Framing
P (%)
P+E (%)
Δ (pp)
Business
62.1
76.9
+14.8
Data science
67.1
76.4
+9.3
Appendix
Table 9: Accuracy by question framing, pooled over all six models and both template families.
Family
Framing
N
P (%)
P+E (%)
Δ (pp)
Data quality
Business
228
86.8
86.8
0.0
Data quality
Data science
228
80.7
81.1
+0.4
Feature contribution
Business
192
32.8
65.1
+32.3
Feature contribution
Data science
192
51.0
70.8
+19.8
Appendix
Table 10: Accuracy by template family and question framing, pooled over all six models.
Condition
Analysis sequence
Final pair
Grade
P
Compare interaction proxies, including empirical mutual information
weather_description , month
Incorrect
P+E
Inspect metadata and target associations; compare six EBM interaction graphs
month, temp
Correct
Appendix
Table 11: GPT-5.6 Luna on the Metro Interstate Traffic Volume business-language dominant-interaction question. The reference pair is month, temp .
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exploratory insight discovery. We introduce DataGovBench, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios. The benchmark includes two tasks: Table QA that requires solving complex decomposable questions and producing textual answers or visualizations, and Table Insight that evaluates the ability of models to generate expert-level findings through exploratory data analysis. Comprehensive experiments with state-of-the-art LLMs, both with and without agentic frameworks, reveal significant performance gaps across both tasks. These results suggest that current LLM-based systems remain far from satisfying the demands of real-world data analytics. DataGovBench provides a challenging benchmark for advancing research on LLMs capable of both answering analytical queries and discovering insights from data. Code and sample data are available at https://github.com/SoHasegawa/datagovbench.
So Hasegawa, Shailaja Keyur Sampat, Lei Liu +1
Fujitsu Research of America, Santa Clara CA 95054, USA
Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two subtasks: (1) Statistical Problem Classification and (2) Variable Identification & Role Assignment. We then introduce StatFormBench, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles. It contains 1,013 samples spanning 20 coarse-grained and 85 fine-grained statistical problem categories. Across 14 open- and closed-source LLMs, the best zero-shot models reach only 72.0 fine-grained classification accuracy and 63.2 variable set overlap. No model performs consistently best across the two subtasks, while enhanced prompting strategies yield only limited or inconsistent gains. We release the benchmark data on Hugging Face at https://huggingface.co/datasets/THU-CongLab/StatFormBench and the evaluation code on GitHub at https://github.com/THU-CongLab/StatFormBench.
Chen Wang, Junzhe Zhao, Xin Cong +2
Department of Statistics and Data Science, Tsinghua University
Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, existing benchmarks remain limited in scope and format. To bridge this gap, we introduce StatABench (Statistical AnalysisBenchmark), a benchmark designed to systematically assess LLMs' statistical analysis capabilities. StatABench comprises two complementary components: Stat-Closed, containing 404 questions across 18 statistical topics in multiple formats (multiple-choice, fill-in-the-blank, decision-making, and practical application), and Stat-Open, featuring 30 complex open-ended modeling tasks adapted from professional competitions. We evaluate diverse LLMs using the LangChain MCP framework and multiple data science agents, and assess Stat-Open solutions via a validated LLM-as-Judge protocol. Experiments show that even GPT-5.1 achieves only 68.6% on Stat-Closed, while the best open-source model reaches 60.6%. On Stat-Open, the top agent framework scores 61.86 on average. These results reveal the gap between current LLMs and reliable statistical analysis, highlighting persistent challenges in tool-grounded reasoning, methodological decision-making, and end-to-end statistical modeling.
Youxin Zhu, Yixuan Ding, Peng Lai +3
Southern University of Science and Technology · The University of Hong Kong · Alibaba Group +1