LabFactory: Building and Evaluating Executable AI Labs
Organizations: University of Oxford
Abstract
Scientific tasks specify a desired capability, but realizing it often requires building a computational system tailored to the task---acquiring data, designing representations, training models, implementing tools, and deciding how they are used at inference. We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface. The builder develops and packages the lab in a metered workspace; a separate host then executes the delivered artifact on held-out inputs, with reference labels kept outside the solver's input interface, and scores its outputs under the task's protocol. This makes the delivered system, rather than the builder's account of its progress, the object of evaluation. We document 28 selected constructions across seven scientific task categories---from molecular and genomic prediction to physiological signals, clinical decision support, and biomedical text---whose delivered labs exceeded their configured reference values on all 33 subtests under host-side execution. Ten contain predictive models fitted during construction; the others assemble retrieval systems, executable analysis environments, and tool-driven workflows around a fixed platform LLM. Together they show that an AI agent can carry a scientific brief all the way to a working lab that can still be invoked, inspected, and checked after construction ends.
Figures & tables
| Task group | Cases | Hours | Tokens M | GPU min | Credits M |
|---|---|---|---|---|---|
| Molecules & proteins | 2 | 19.54 | 23.86 | 544.5 | 1.04 |
| Genomics & omics | 4 | 12.59 | 39.16 | 299.4 | 2.10 |
| Imaging | 1 | 20.09 | 28.80 | 217.4 | 0.49 |
| Physiological signals | 3 | 18.72 | 59.38 | 957.7 | 0.45 |
| Knowledge & reasoning | 8 | 15.39 | 15.45 | 46.9 | 4.37 |
| Clinical decision support | 2 | 17.01 | 36.11 | 9.5 | 9.98 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Subtest | Metric | Solver | Reference bar | Reference source | |
|---|---|---|---|---|---|---|
| Protein–protein interaction | Bernett PPI | AUROC | 3000 | 0.9874 | 0.7000 | Liu et al.,2025 |
| Protein–ligand binding affinity | PDBbind | Pearson | 285 | 0.9141 | 0.8030 | Graber et al., 2025 7 |
| Causal eQTL variants | eQTL | AUPRC | 8862 | 0.9192 | 0.5960 | Wang et al., 2025 § |
| Bioinformatics data analysis | BixBench | Judge acc. | 296 | 0.4561 | 0.2280 | Ghareeb et al., 2026 1 |
| 5mC methylation sites | 5mC | AUROC | 1407 | 0.9530 | 0.7830 | Feng et al.,2025 |
| Polyadenylation signal detection | Poly(A) signal | F1 | 3000 | 0.8977 | 0.8800 | Shen et al.,2026 |
| Subtest | Raw LLM | Solver | (pp) | |
|---|---|---|---|---|
| MedMCQA | 4183 | 0.7930 | 0.8735 | +8.05 |
| MEDIQA | 1107 | 0.7868 | 0.8193 | +3.25 |
| MetaMedQA | 1373 | 0.8267 | 0.8529 | +2.62 |
| LabSafety | 632 | 0.8481 | 0.8718 | +2.37 |
| MEDEC | 597 | 0.8375 | 0.8526 | +1.51 |
| Medbullets | 308 | 0.8994 | 0.9123 | +1.29 |