AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
Organizations: Department of Computer Science and Engineering · The Ohio State University · Department of Chemistry · University of Wisconsin–Madison · Department of Psychology · Department of Geography · College of Pharmacy · Department of Biomedical Informatics · Cisco Research
Abstract
Despite long-standing efforts in accelerating scientific discovery with AI, building AI co-scientists remains challenging due to limited high-quality data for training and evaluation. To tackle this data scarcity issue, we present AutoSDT, an automatic pipeline that collects high-quality coding tasks in real-world data-driven discovery workflows. AutoSDT leverages the coding capabilities and parametric knowledge of LLMs to search for diverse sources, select ecologically valid tasks, and synthesize accurate task instructions and code solutions. Using our pipeline, we construct AutoSDT-5K, a dataset of 5,404 coding tasks for data-driven discovery that covers four scientific disciplines and 756 unique Python packages. To the best of our knowledge, AutoSDT-5K is the only automatically collected and the largest open dataset for data-driven scientific discovery. Expert feedback on a subset of 256 tasks shows the effectiveness of AutoSDT: 93% of the collected tasks are ecologically valid, and 92.2% of the synthesized programs are functionally correct. Trained on AutoSDT-5K, the Qwen2.5-Coder-Instruct LLM series, dubbed AutoSDT-Coder, show substantial improvement on two challenging data-driven discovery benchmarks, ScienceAgentBench and DiscoveryBench. Most notably, AutoSDT-Coder-32B reaches the same level of performance as GPT-4o on ScienceAgentBench with a success rate of 7.8%, doubling the performance of its base model. On DiscoveryBench, it lifts the hypothesis matching score to 8.1, bringing a 17.4% relative improvement and closing the gap between open-weight models and GPT-4o.
Figures & tables
| Statistics | Value |
| # Tasks | 5,404 |
| # Repositories | 1,325 |
| # Packages | 756 |
| Cost (USD) | 2,955 |
| Disciplines (# Tasks/ # Repositories): | |
| Bioinformatics | 1,466 / 396 |
| Difficulty | % | Avg # of Lines | Avg # of Subtasks |
| Easy | 22.3 | 214.7 | 4.1 |
| Medium | 48.4 | 263.7 | 4.4 |
| Hard | 29.3 | 403.2 | 5.1 |
| Model Size | ScienceAgentBench | DiscoveryBench | |||||||
| SR(%, ) | VER (%, ) | HMS (%, ) | |||||||
| Base | SFT | Base | SFT | Base | SFT | ||||
| 7B | 3.3 ( 0.5) | 2.3 ( 0.9) | -1.0 (30%) | 19.9 ( 0.2) | 27.5 ( 3.3) | +7.6 (38%) | 4.8 ( 1.0) | 6.3 ( 1.3) | +1.5 (31%) |
| 14B | 4.3 ( 0.5) | 5.9 ( 1.6) | +1.6 (37%) | 26.5 ( 2.1) | 35.0 ( 2.5) | +8.5 (32%) | 6.4 ( 0.2) | 7.3 ( 0.3) | +0.9 (14%) |
| 32B | 3.9 ( 0.8) | 7.8 ( 1.4) | +3.9 (100%) | 28.4 ( 0.8) | 36.0 ( 5.3) | +7.6 (27%) | 6.9 ( 0.6) | 8.1 ( 0.7) | +1.2 (17%) |
| Models | SR (%, ) | VER (%, ) |
| Proprietary Reasoning Models | ||
| Claude-3.7-Sonnet | 18.6 ( 0.8) | 51.6 ( 4.7) |
| OpenAI o1-preview | 23.9 ( 0.5) | 56.2 ( 1.7) |
| Proprietary Non-Reasoning Models | ||
| GPT-4o (2024-05-13) | 7.5 ( 0.5) | 42.2 ( 1.6) |
| GPT-4o (2024-11-20) | 11.4 ( 1.2) | 43.1 ( 2.1) |
| Training Data | Bio. | Chem. | Geo. | Psy & Neu |
| Bio-only | 18.5 | 10.0 | 0.0 | 7.1 |
| Chem-only | 11.1 | 15.0 | 0.0 | 7.1 |
| Geo-only | 14.8 | 15.0 | 3.7 | 7.1 |
| Psy & Neu | 11.1 | 5.0 | 3.7 | 7.1 |
| Full | 11.1 | 15.0 | 14.8 | 7.1 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Task | Subject | Scientific | Naturally | Auto |
| Instances | Domains | Dataset | Occurring Code | Collection | |
| DA-Code ( Huang et al., 2024 ) | 500 | 0 | ✗ | ✓ | ✗ |
| DSBench ( Jing et al., 2025 ) | 540 | 0 | ✗ | ✗ | ✗ |
| MLE Bench ( Chan et al., 2025 ) | 75 | 1 | ✗ | ✗ | ✗ |
| REBench ( Wijk et al., 2024 ) | 7 | 1 | ✗ | ✗ | ✗ |
| ScienceAgentBench ( Chen et al., 2025 ) | 102 | 4 | ✓ | ✓ | ✗ |
| You are an expert at reading GitHub README.md files thoroughly and determining whether the repository hosts code related to a research paper or not, and you are also skilled at correctly extracting the link to the related paper. Your answer should be based on your thorough understanding of the content of the README.md file. Does the README.md file indicate that the repository hosts code related to a research paper in the discipline of {keyword} ? Answer by ‘YES’ or ‘NO’ in the ‘RESEARCH’. If your answer to the previous question is ‘YES’, extract the link to the related research paper. Make sure to extract the link to the research paper that this repository implements only, this should be the link to the paper that people would cite if they used the code in the repository for their work, ignoring all other irrelevant links that might be referenced in the README.md file. Put the link(s) in front of the ‘LINKS’: as a list of links. README.md file: {readme} You should strictly follow the format below: RESEARCH: LINKS: |
| You are an expert at determining whether a program contains scientific code or not. Given a code file, you need to verify if the current code is a scientific task. Several conditions should be satisfied: 1. Functionality: the functionality of the given program should be related to tasks in a scientific workflow. These tasks include but are not limited to feature engineering, machine learning, deep learning, computational analysis, data visualization, model training, numerical calculation/analysis, statistical methods, domain-specific analysis/simulation, etc. 2. Input: the program should receive at least one or multiple datasets as input. In other words, the program is dealing with a dataset and conducting analysis or experiments on top of the data. The data can either be loaded through built-in functions or be loaded from local files. If the current program does not receive and process any data, it cannot be considered as “a scientific task” here. 3. Output: the program should output numerical or visualization results that can be further evaluated. A code file is considered a scientific task ONLY IF it completely satisfied the three dimensions above. For example, code files that purely contain modeling, training/testing, data pre-processing, or only consist of utility functions or class definitions, are not considered a scientific task. Program name: {file_name} Program code: {code} After reasoning about the problem, output your final answer strictly based on the following format: VERDICT: {YES/NO} |
| You are an expert software engineer who is very skilled at analyzing Python code files and their repositories to extract dependencies. In this task you will be given a Python file and the GitHub file tree of the repository it belongs to, your job is to thoroughly understand the code and all the in-repository dependencies it needs. This is because we would like to run this code in a standalone environment and we have to make sure that all the dependencies that the code needs are copied in that environment. Hence, it is very important that you have a thorough understanding of the code and extract all in-repository dependencies needed. Specifically, your job is to do the following: 1. Recognize whether the code makes use of a dataset. The dataset can either be loaded via built-in library functions (e.g., data = MNIST ()) or loaded from a local file in the repository (csv, jsonl, xls, txt, parquet, or any other file type). If the dataset(s) used in the code are either loaded through built-in library functions or contained within the repository, you should output “Yes” in DATASET_LABEL field. Otherwise, you should output “No”. 2. In the case where the dataset used in the code is contained within the repository, you also have to find the relative path to the dataset file, based on the GitHub file tree that will be given to you. You will list the paths to all datasets used in the code as a list of paths after the field DATASET_PATHS. 3. Besides the dataset, now you have to identify all other in-repository dependencies that the code uses, and extract their relative paths based on the file tree given to you. These can be modules, classes, models, or any other dependency that the code imports from a folder within the repository. If you identify that there are in-repository dependencies used, you should put a “Yes” in the MODULE_LABEL. Otherwise, output a “No”. 4. In the case of a “Yes”, make sure to put the relative paths to all dependencies as a list of paths in the MODULE_PATHS field, based on the GitHub file tree given to you. 5. If based on the code alone you can only identify the folder that contains the dependency but not the exact file only return the path to the folder. This is because you might sometimes not be able to know which file the dependency is exactly located in based on only looking at the file tree. Thus, to stay on the safe side, just give the path to the folder that contains the dependency. Python code: {code} Project directory: {directory} |
| You are an excellent coder at adapting existing files for standalone executability. You will be given a code file from a Github Repo. Your task is to modify the code into a self-contained program that can be run locally and separately. Please do not change the original functionality of the code. You must keep the original logic and functionality of the code as much as possible. You should never include dummy/pass statements or empty/mock functions in your response. You need to slightly modify the source code’s input/output logistics and intermediate steps to make it a stand-alone program that can be executed locally. The modified code will then be executed in a local environment. If there are errors, you need to debug the code based on the execution feedback. All the datasets and dependency files are located at {dataset_path} . If the original code has imported modules from local files, you can assume they exist and do the same imports in your modified code. Here is the directory structure of the dataset and dependency files: {dataset_structure} Make sure that the code you generate uses the same input files as the original code. Do not generate dummy input files or input data. For the output of the programs, your code should save the results to a file named ‘‘pred_results/pred_[code_file_name].[extension]’’ , depending on the type of data such as csv , txt , jsonl , etc. ALL outputs of the program should be saved in the directory pred_results/ . You should never create new folders or files outside of the specified directory. Code to be modified: {code_file_name} {code} The user may execute your code and report any exceptions and error messages. You should address the reported issues and respond with a fixed, complete program. Note that, when addressing bugs, you should ONLY focus on addressing the errors and exceptions and MUST NOT change or delete the main functionality and logic of the original program just to make it executable. Keep your response concise and do not use a code block if it’s not intended to be executed. Do not suggest a few line changes, incomplete program outline, or partial code that requires the user to modify. Your response should include a complete, standalone, executable program. Do not use any interactive Python commands in your program, such as ‘!pip install numpy‘, which will cause execution errors. Regardless of the iterations of self-debugging, make sure to wrap your program in a code block that specifies the script type, python. For example: “‘python print("Hello World!") ”’ |
| You are a helpful agent for generating task instructions based on a code snippet for solving scientific data processing tasks. You need to provide a clear and concise instruction that best describes the functionality of the given code. The instruction should be written in plain English and should be detailed enough so that a person who has no knowledge of the code can understand the task and implement code for it. The instruction should not reveal too many implementation details but also should be precise and not vague. It should be a high-level description of the code’s functionality. You should thoroughly read the scientific data processing code snippet provided, understand the underlying domain-specific concepts behind it, and generate a task instruction that makes correct use of the domain-specific language. In other words, your task instructions should be written as if they are from a domain scientist giving instructions to a junior researcher in their lab. The structure of the instruction should be as clear as possible: you should clearly specify the goal of the task, clearly name the exact input file/files that should be used, and the output files that should be created and the path to which they should be saved. Additionally, if the output of the program is written to a file, you should specify the format that the output should be written in, based on the implementation given in the code snippet. In cases where, based on your understanding of the code, you deem that the instruction needs more details - for example, if a certain program can use different computational methods to reach a solution - you can add guidelines about the specific method to use in the instruction. In all cases, ensure that the instruction does not include too many implementation details but also that it is precise and does not invite ambiguity or confusion. The format of your instruction should be a concise paragraph of a few lines without any sections. Keep the instruction focused on the high level scientific goal of the task and do not make reference to unnecessary details like "ensure the directory or so and so files exist". Such low level implementation details should never be part of the instruction. Please generate the instruction based on the code snippet below. {code} |
| Stage | Cost (USD) |
| AutoSDT-Search: | |
| Repository Crawling | 32 |
| AutoSDT-Select: | |
| Scientific Task Filtering | 459 |
| Dependency Locating | 828 |
| AutoSDT-Adapt: |
| Model Size | HMS(%) |
| Re-implementation Results | |
| Qwen2.5-Coder-7B-Instruct | 4.8 |
| AutoSDT-Coder-7B | 6.3 |
| Qwen2.5-Coder-14B-Instruct | 6.4 |
| AutoSDT-Coder-14B | 7.3 |
| Qwen2.5-Coder-32B-Instruct | 6.9 |
| Discipline | Task Instruction |
| Bioinformatics | Predict circRNA-disease associations using the Random Walk with Restart (RWR) algorithm. Utilize the circRNA-disease association data in “circrna_disease.txt”, along with circRNA and disease lists from “circ_list.csv” and “dis_list.csv” respectively. Perform 5-fold cross-validation to evaluate prediction performance, calculating metrics such as accuracy, recall, precision, F1-score, AUC, and AUPR. Save the results to “RWR.csv” in CSV format, including metrics and their values. |
| Computational Chemistry | Cluster molecular structures based on their chemical fingerprints using the SMILES data in “smiles.csv”. Compute Morgan fingerprints for the molecules, perform clustering using the Butina algorithm with a similarity cutoff of 0.72, and identify the centroid molecule for each cluster. Save the clustering summary, including the number of clusters and centroid SMILES, to “clustering.txt” and generate SVG visualizations of the centroid molecules for each cluster and save them as “centroid.svg”. |
| Geographic Inf. Sci. | Match geo-tagged drone images to corresponding satellite map images using geographic coordinates. Use the satellite map data from “map.csv” and the drone photo metadata from “metadata.csv”. For each drone image, determine its location on the satellite map by comparing its geographic coordinates with the boundaries of the satellite images. Calculate the drone image’s precise geographic position within the matched satellite image and compare it to the ground truth coordinates. Save the results, including the calculated coordinates, errors, and matching status, to “results.csv”. |
| Psy. and Cog. Neuroscience | Process MRI data to calculate the incidence sizes of parental brain regions. Use the MRI in mri.nii.gz and the Allen Brain annotation file allen.nii.gz. Apply a threshold to identify stroke-affected regions and generate the following outputs: (1) a labeled NIfTI file highlighting affected regions saved as “affected_regions_parental.nii.gz”, (2) a text file summarizing stroke volume and affected region percentages saved as “summary.txt”, and (3) a MATLAB file with detailed region labels and metrics saved as “label_count.mat”. |
| License | Repositories |
| MIT | 449 |
| GNU | 247 |
| Apache | 145 |
| BSD | 84 |
| CC | 57 |
| Boost | 4 |
| Repositories |
| GabrieleLozupone/AXIAL |
| fhalab/MLDE |
| snacktavish/TreeToReads |
| usnistgov/SDNist |
| ruppinlab/CSI-Microbes-identification |
| fenchri/edge-oriented-graph |