cs.LGOct 8, 2026

DataSense-Bench: The First Step Toward an AI Scientist

Authors: Yudi Zhang, Mingyu Cao, Lu Yin, Mykola Pechenizkiy, Shiwei Liu

Organizations: Eindhoven University of Technology · Max Planck Institute for Intelligent Systems · University of Surrey · ELLIS Institute Tübingen · Tübingen AI Center

Abstract

As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training? We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning. We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model. Agents are allowed to inspect the data, write and execute analysis code, and run model forward passes, but can not train the model or access the actual evaluation tasks. We then fine-tune the base model on each selected subset and evaluate its post-training performance under a standardized protocol. We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively. We then evaluate the agents along two complementary dimensions: the post-training performance of the top-ranked subset, reflecting the ability to identify high-value training data, and ranking accuracy, reflecting the ability to predict the relative performance of the selected subsets. In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs. Analysis of execution traces on both tasks shows that agents often use similar data signals while interpreting their training value differently.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

    Jul 28, 2026Fanqing Meng, Lingxiao Du, Qiguang Chen +4LLM Agent EvaluationLLM Agent Self-Improvement

  2. AutoData: Agentic Search for Pre-training Data Selection

    Sep 17, 2026Yan Meng, Dhruv Srikanth, Bingchen Zhao +2Language Model PretrainingAutomated Algorithm Discovery

  3. What is Missing from AI Post-Training AI: An Empirical Analysis

    Aug 19, 2026Joy Jia Yin Lim, Xin Huang, Hao Peng +5AI Agent EvaluationLLM Agent Self-Improvement