cs.LGSep 28, 2026

SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals

Authors: Jingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez, Joshua Shunk, Rich Caruana, Ben Lengerich

Organizations: University of Wisconsin–Madison, Intelligible · Intelligible · University of Cincinnati

Abstract

Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

    Jul 7, 2026So Hasegawa, Shailaja Keyur Sampat, Lei Liu +1Large Language Model BenchmarksBenchmark Datasets

  2. Benchmarking Language Models for Statistical Problem Formulation

    Sep 2, 2026Chen Wang, Junzhe Zhao, Xin Cong +2Large Language Model BenchmarksStatistical Inference

  3. StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs

    Jun 22, 2026Youxin Zhu, Yixuan Ding, Peng Lai +3Statistical InferenceTool-Grounded Reasoning