Science sandboxes measure the scientific capability of AI agents
Organizations: The Broad Institute of MIT and Harvard; Cambridge, MA 02142, USA. · Department of Biomedical Informatics, Harvard Medical School; Boston, MA 02115, USA. · The Jackson Laboratory; Bar Harbor, ME 04609, USA. · Sutter Hill Ventures; Palo Alto, CA 94304, USA. · Department of Systems Biology, Harvard Medical School, Boston, MA, USA. · David H. Koch Institute for Integrative Cancer Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA. · Howard Hughes Medical Institute, Chevy Chase, MD 20815, USA. · The Wyss Institute for Biologically Inspired Engineering at Harvard University, Boston, MA 02115, USA. · Harvard-MIT Program in Health Sciences and Technology, Institute for Medical Engineering and Science, Massachusetts Institute of Technology, Cambridge, MA 02139, USA. · Department of Genetics, Yale School of Medicine; New Haven, CT, USA. · Department of Biology, Massachusetts Institute of Technology, Cambridge, MA, USA. · Department of Immunology and Infectious Diseases, Harvard T.H. Chan School of Public Health, Boston, MA 02115, USA. · Department of Organismic and Evolutionary Biology, Harvard University, Cambridge, MA 02138, USA.
Abstract
Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI agents through repeated cycles of experimentation, feedback, and hypothesis revision. Science sandboxes invite an agent to query the natural world in different ways, ranging from "wet" physical experiments, to "damp" predictive models trained on empirical data, to "dry" invented rules. By establishing a common experimental loop and a protocol for evaluating agents within it, science sandboxes allow assessment of both quantitative performance on specific metrics and qualitative scientific reasoning, across a spectrum of empirical verifiability. Here, we instantiate this framework in two biological settings, models of regulatory genomics and protein fitness prediction, and examine the capabilities of frontier agents. Across these settings, we could see when agents successfully optimized a quantitative metric without understanding the rules underlying the system. In particular, their scientific reasoning deteriorated when they encountered systems whose rules fell outside familiar biological priors. By highlighting such failure modes, science sandboxes make the frontier of scientific capability measurable and provide a controlled setting in which to study and ultimately expand it.