SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals
Organizations: University of Wisconsin–Madison, Intelligible · Intelligible · University of Cincinnati
Abstract
Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.
Figures & tables
| Family | Templates | Representative injections | Representative tasks |
| Data quality | 8 | Corrupt a subset of target values and add a binary column marking those rows; replace some category values with a missing-like label; add a group column with one very rare group. | Identify the marker column or the missing-like label; decide whether there is sufficient statistical evidence to make claims about the rare group. |
| Feature contribution | 9 | Permute one numerical feature to break its link to the target; add a smooth inverted-U effect of one feature on the target; add an XOR-shaped effect on the target over two numerical features. | Name the noise feature; name the feature and the value at which its effect peaks; name the interacting pair or its conditional direction. |
| Framing | Data quality | Feature contribution | Total |
| Data science | 38 | 32 | 70 |
| Business | 38 | 32 | 70 |
| Total | 76 | 64 | 140 |
| Model | Data quality | Feature contribution | Total |
| Claude Sonnet 5 | 77.6 | 18.8 | 50.7 |
| Claude Opus 4.8 | 84.2 | 29.7 | 59.3 |
| Claude Fable 5 | 85.5 | 45.3 | 67.1 |
| GPT-5.6 Sol | 86.8 | 71.9 | 80.0 |
| GPT-5.6 Terra | 85.5 | 51.6 | 70.0 |
| GPT-5.6 Luna | 82.9 | 34.4 | 60.7 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Template | What the injection changes | Question wording (data-science and business) |
|---|---|---|---|
| Data quality | Bad-row indicator | Corrupt a subset of target values and add a binary column whose positive entries mark exactly those rows. | Data-science: You can apply exactly one rule based on a single column to exclude problematic rows before analyzing {OUTCOME_COL} . Which column should the rule be based on? Return a single column name. Business: Some records have already been marked for review in a yes-or-no column. Which column contains that review marker? Return a single column name. |
| Conditional anomaly indicator | Rewrite 5% of outcomes so that they remain within the observed support but become surprising given similar records. Add the true binary flag and same-prevalence placebo flags. | Data-science: Among the binary 0/1 columns, one indicator has 1 exactly on rows whose {OUTCOME_COL} outcome is conditionally inconsistent with otherwise similar records. The injected outcomes remain within the original target support. Which column is the indicator? Return a single column name. Business: Among the 0/1 columns, one indicator’s 1 values mark records whose {OUTCOME_COL} outcome is inconsistent with otherwise similar records, even though the outcome itself remains within the dataset’s observed range. Which column is that indicator? Return a single column name. | |
| Category target outlier | Choose a sufficiently large category and shift its target values up or down until its mean stands out from the other categories. | Data-science: Is there a category in {FEATURE} whose {OUTCOME_COL} values are unusually high or low compared with the other categories? Return only the category label. Business: Is there a category in {FEATURE} whose {OUTCOME_COL} values stand out from the other categories? Return only the category label. | |
| Conditional bad rows | Change the targets of a small set of rows so they remain plausible in the overall target distribution but are unusual given their feature values. | Data-science: Some rows have {OUTCOME_COL} values that are plausible globally but unusual given their feature values. Identify the {N_INJECT_SAMPLES} affected rows and return their row_id values as a comma-separated list. Business: There are {N_INJECT_SAMPLES} rows in this data with an unusual pattern in {OUTCOME_COL} . Return the row_id values as a comma-separated list. | |
| Underpowered group evidence | Add a categorical group column with two categories: a common group A and a very rare group B; leave the target unchanged so the group comparison has a wide interval. | Data-science: For categorical feature {GROUP_COL} , is there enough statistical evidence to distinguish group A from group B with respect to {OUTCOME_COL} ? Return one of: enough evidence, not enough evidence. Business: For the categorical feature {GROUP_COL} , is the evidence strong enough to say the groups behave differently with respect to {OUTCOME_COL} ? Return one of: enough evidence, not enough evidence. | |
| Semantic missing label | Replace a small number of values in a categorical feature with a distinct missing-like label such as “Unknown” or “Not reported.” | Data-science: Which category label in {FEATURE} appears to encode missing, unknown, or not-reported values? Return only the label. Business: Are there any labels in {FEATURE} that we need to pay attention to? If yes, return only the label. If no, return No. |
| Dataset | Source rows | Target | Reference |
| AI4I 2020 Predictive Maintenance | 10,000 | Machine failure | ( AI4I, 2020 ) |
| Bike Sharing | 17,379 | Rental count | ( Fanaee-T, 2013 ) |
| California Housing | 20,640 | House value | ( scikit-learn, n.d. ) |
| King County Housing | 21,613 | Sale price | ( King County Housing, n.d. ) |
| Metro Interstate Traffic Volume | 48,204 | Traffic volume | ( Hogue, 2019 ) |
| Steel Industry Energy Consumption | 35,040 | Energy use | ( Sathishkumar V E et al., 2021 ) |
| Data quality | Feature contribution | Total | |||||||
| Model | P | P+E | P | P+E | P | P+E | |||
| Claude Sonnet 5 | 77.6 | 77.6 | 0.0 | 18.8 | 53.1 | +34.4 | 50.7 | 66.4 | +15.7 |
| Claude Opus 4.8 | 84.2 | 84.2 | 0.0 | 29.7 | 65.6 | +35.9 | 59.3 | 75.7 | +16.4 |
| Claude Fable 5 | 85.5 | 88.2 | +2.6 | 45.3 | 75.0 | +29.7 | 67.1 | 82.1 | +15.0 |
| GPT-5.6 Sol | 86.8 | 85.5 | -1.3 | 71.9 | 78.1 | +6.3 | 80.0 | 82.1 | +2.1 |
| GPT-5.6 Terra | 85.5 | 85.5 | 0.0 | 51.6 | 82.8 | +31.3 | 70.0 | 84.3 | +14.3 |
| Model or family | P: C/P/I | P+E: C/P/I | |
| Claude Sonnet 5 | 140 | 71/12/57 | 93/8/39 |
| Claude Opus 4.8 | 140 | 83/9/48 | 106/8/26 |
| Claude Fable 5 | 140 | 94/8/38 | 115/10/15 |
| GPT-5.6 Sol | 140 | 112/9/19 | 115/9/16 |
| GPT-5.6 Terra | 140 | 98/9/33 | 118/6/16 |
| GPT-5.6 Luna | 140 | 85/8/47 | 97/10/33 |
| Model | Framing | P (%) | P+E (%) | (pp) | |
| Claude Sonnet 5 | Business | 70 | 47.1 | 67.1 | +20.0 |
| Claude Sonnet 5 | Data science | 70 | 54.3 | 65.7 | +11.4 |
| Claude Opus 4.8 | Business | 70 | 52.9 | 74.3 | +21.4 |
| Claude Opus 4.8 | Data science | 70 | 65.7 | 77.1 | +11.4 |
| Claude Fable 5 | Business | 70 | 62.9 | 84.3 | +21.4 |
| Claude Fable 5 | Data science | 70 | 71.4 | 80.0 | +8.6 |
| Framing | P (%) | P+E (%) | (pp) |
| Business | 62.1 | 76.9 | +14.8 |
| Data science | 67.1 | 76.4 | +9.3 |
| Family | Framing | P (%) | P+E (%) | (pp) | |
| Data quality | Business | 228 | 86.8 | 86.8 | 0.0 |
| Data quality | Data science | 228 | 80.7 | 81.1 | +0.4 |
| Feature contribution | Business | 192 | 32.8 | 65.1 | +32.3 |
| Feature contribution | Data science | 192 | 51.0 | 70.8 | +19.8 |
| Condition | Analysis sequence | Final pair | Grade |
| P | Compare interaction proxies, including empirical mutual information | weather_description , month | Incorrect |
| P+E | Inspect metadata and target associations; compare six EBM interaction graphs | month, temp | Correct |