DataSense-Bench: The First Step Toward an AI Scientist
Organizations: Eindhoven University of Technology · Max Planck Institute for Intelligent Systems · University of Surrey · ELLIS Institute Tübingen · Tübingen AI Center
Abstract
As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training? We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning. We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model. Agents are allowed to inspect the data, write and execute analysis code, and run model forward passes, but can not train the model or access the actual evaluation tasks. We then fine-tune the base model on each selected subset and evaluate its post-training performance under a standardized protocol. We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively. We then evaluate the agents along two complementary dimensions: the post-training performance of the top-ranked subset, reflecting the ability to identify high-value training data, and ranking accuracy, reflecting the ability to predict the relative performance of the selected subsets. In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs. Analysis of execution traces on both tasks shows that agents often use similar data signals while interpreting their training value differently.
Figures & tables
| Rank | Agent / reference | gain | (%) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| BFCL accuracy (50 trajectories per group) Base model: 31.67% | |||||||||
| 1 | DeepSeek V4 Pro | 36.53_{\text{\pm1.83}} [-2pt] Max: 38.62 | 1.78 | \mathbf{36.61}_{\text{\pm0.72}} [-2pt]Max: 37.42 | 35.74_{\text{\pm1.12}} [-2pt]Max: 37.00 | 36.08_{\text{\pm1.44}} [-2pt]Max: 37.75 | 34.33_{\text{\pm1.16}} [-2pt]Max: 35.67 | 0.65 | 33.33 |
| 2 | GPT-6 Astra | \mathbf{36.49}_{\text{\pm0.23}} [-2pt]Max: 36.75 | 1.74 | 35.31_{\text{\pm0.94}} [-2pt]Max: 36.38 | 34.74_{\text{\pm1.50}} [-2pt]Max: 36.21 | 34.93_{\text{\pm0.42}} [-2pt]Max: 35.42 | 34.58_{\text{\pm0.07}} [-2pt]Max: 34.62 | 0.83 | 100.00 |
| 3 | Kimi K3 | 36.29_{\text{\pm0.91}} [-2pt]Max: 36.92 | 1.54 | 36.21_{\text{\pm0.65}} [-2pt]Max: 36.92 | 36.18_{\text{\pm1.22}} [-2pt]Max: 37.58 | 35.85_{\text{\pm1.14}} [-2pt]Max: 36.62 | \mathbf{36.46}_{\text{\pm1.27}} [-2pt]Max: 37.75 | 0.02 | 33.33 |
| 4 | Opus 5 | 35.94_{\text{\pm0.72}} [-2pt]Max: 36.75 | 1.19 | \mathbf{36.35}_{\text{\pm1.09}} [-2pt]Max: 37.38 | 35.35_{\text{\pm0.53}} [-2pt]Max: 35.96 | 34.29_{\text{\pm2.29}} [-2pt]Max: 36.54 | 31.08_{\text{\pm3.45}} [-2pt]Max: 35.04 | 0.77 | 0.00 |
| 5 | Fable 5.1 | 35.21_{\text{\pm0.94}} [-2pt]Max: 36.00 | 0.46 | 34.96_{\text{\pm1.67}} [-2pt]Max: 36.58 | 34.87_{\text{\pm0.07}} [-2pt]Max: 34.96 | 34.10_{\text{\pm0.24}} [-2pt]Max: 34.38 | \mathbf{35.57}_{\text{\pm1.53}} [-2pt]Max: 37.17 | 0.23 | 33.33 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Terminal problem solving | Multi-turn tool use | |
| Training pool | OpenThoughts-Agent 10K | EnvScaler SFT |
| Groups group size | ||
| Base model | Qwen3-4B | Qwen3-4B |
| Selectors / runs per selector | 10 / 3 | 6 / 3 |
| Selection GPU / time budget | H100 80GB / 4 h | H100 80GB / 4 h |
| Learning rate / epochs | / 7 | / 2 |
| Agent | Model identifier |
|---|---|
| Fable 5.1 | claude-fable-5-1 |
| Opus 5 | claude-opus-5 |
| GPT-6 Astra | gpt-6-astra |
| Kimi K3 | kimi-k3 |
| DeepSeek V4 Pro | deepseek-v4-pro |
| Gemini 3.5 Flash | gemini-3.5-flash |
| Agent | Run | Signals used | Final grouping rule |
|---|---|---|---|
| Fable 5 | 1 | Structure, NLL tails and 64 task clusters. | Cluster-balanced ; spaced lower score bands. |
| 2 | Structure and partial NLL coverage. | Task deduplication and cluster caps in . | |
| 3 | Structure and 8K-prefix NLL. | Task-deduplicated without clustering; lower score windows. | |
| Fable 5.1 | 1 | Structural quality with a small JSON-NLL term. | Task deduplication and cluster/style caps in / ; lower score bands. |
| 2 | Structural quality, model loss and a prompted concreteness score. | Task deduplication and cluster/style caps in / ; lower score bands. | |
| 3 | Structural quality with soft loss-anomaly penalties. | Cluster/source caps and an 8% Tezos cap in / ; next-ranked / and bottom . |
| Agent | Recipe and cross-run differences |
|---|---|
| Fable 5.1 | Run 1 combines schema checks, errors, reasoning length, and length-adjusted loss; it samples descending score bands, ending with low-scoring rows. Run 2 emphasizes reasoning–action consistency and loss outliers. Run 3 separately scores reasoning, actions given reasoning, and actions without reasoning, then combines quality gates with environment coverage. Later groups include increasingly defect-heavy trajectories. |
| Opus 5 | Run 1 prioritizes structural quality and uses a small residualized-loss term, with fixed interaction-type and environment coverage across groups. Run 2 adds verbosity-normalized reasoning loss and recovered-error bonuses, placing noncanonical argument formats in the lower groups. Run 3 probes action difficulty without the demonstrated reasoning and selects widely separated score bands with distinct environments. |
| GPT-6 Astra | All runs combine schema checks, argument grounding, state-changing actions, clarification, and coverage. Run 1 adds sampled-turn loss over the pool and inspection of top candidates. Run 2 uses a bounded preference for moderate reasoning loss and a fixed 70/30 interaction-type mix. Run 3 scores a quality-filtered shortlist and adds semantic checks on unsupported actions. Groups occupy separated predicted-value bands. |
| Kimi K3 | Run 1 uses loss to penalize outliers, normalizes quality within interaction types, and selects successive groups under diversity caps. Run 2 favors intermediate total and tool-call loss. Run 3 instead rewards lower loss alongside successful responses and tool coverage, with environment caps relaxed from one to three across the groups. |
| DeepSeek V4 Pro | Run 1 prioritizes high sampled-turn loss, trajectory length, and final-action loss. Run 2 combines loss with structural richness, including a positive failed-call term, and uses diversity-penalized greedy selection. Run 3 emphasizes tool breadth, penalizes failed-call rate, and adds within-environment loss as a secondary signal. |
| Gemini 3.5 Flash | Run 1 separates groups by observed tool success, length, and turn count. Run 2 uses a heuristic score penalizing errors and loops while rewarding reasoning length and intermediate turn counts. Run 3 adds base-model loss to choose long, error-free trajectories for the top groups; short clean trajectories and error-heavy trajectories occupy lower groups. |
| Agent | Run | rank | ||||||
|---|---|---|---|---|---|---|---|---|
| Fable 5 | r1 | 0.109 | 0.085 | 0.145 | 0.101 | 0.107 | 0.10 | 2 |
| r2 | 0.100 | 0.110 | 0.121 | 0.110 | 0.099 | 0.30 | 4 | |
| r3 | 0.121 | 0.121 | 0.135 | 0.134 | 0.121 | -0.50 | 4 | |
| Fable 5.1 | r1 | 0.117 | 0.122 | 0.129 | 0.098 | 0.116 | 0.50 | 3 |
| r2 | 0.110 | 0.103 | 0.110 | 0.097 | 0.132 | -0.30 | 3 | |
| r3 | 0.148 | 0.111 | 0.121 | 0.112 | 0.101 | 0.70 | 1 |
| Agent | Run | rank | ||||||
|---|---|---|---|---|---|---|---|---|
| Fable 5.1 | r1 | 36.00 | 33.25 | 34.83 | 34.38 | 37.17 | -0.30 | 2 |
| r2 | 34.17 | 36.58 | 34.96 | 33.96 | 34.12 | 0.60 | 3 | |
| r3 | 35.46 | 35.04 | 34.83 | 33.96 | 35.42 | 0.40 | 1 | |
| Opus 5 | r1 | 35.71 | 36.46 | 35.08 | 31.96 | 28.75 | 0.90 | 2 |
| r2 | 36.75 | 37.38 | 35.00 | 36.54 | 29.46 | 0.80 | 2 | |
| r3 | 35.38 | 35.21 | 35.96 | 34.38 | 35.04 | 0.60 | 2 |