q-bio.OTJun 27, 2026

Building AI-Ready Data Systems for Space Life Sciences, Aerospace Medicine, and Deep Space Exploration

Authors: Sylvain V. CostesSergio Garcia BustoRyan T. ScottJames A. CasalettoGautier Bardi de FourtouBrian M. EvartsAmanda M. Saravia-ButlerXavier-Lewis Palmer+10 more

Organizations: Blue Marble Space, Seattle, Washington, USA · Computational and Systems Biology, University of Pittsburgh, Pittsburgh, PA, USA · Trivedi Institute for Space and Global Biomedicine, University of Pittsburgh, Pittsburgh, PA, USA · Wellcome Sanger Institute, Wellcome Genome Campus, Hinxton, CB10 1SA, UK · University of Cambridge, Cambridge, UK · Department of Systems Biology, Harvard Medical School, Boston, MA, USA · Amentum, Space Biosciences Division, NASA Ames Research Center, Moffett Field, CA, USA · Massachusetts Institute of Technology, Cambridge, USA · Crown Point Technologies, Inc, Columbia, MD, 21046 · BiosView Labs, Dayton, Ohio, USA · Directorate of Human and Robotic Exploration, European Space Agency, Noordwijk, the Netherlands · Texas State University, USA · McGowan Institute for Regenerative Medicine, Department of Surgery, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA · Center for Space Biomedicine, Trivedi Institute for Space and Global Biomedicine, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA · Department of Bioengineering, University of Pittsburgh, Pittsburgh, PA, USA · Stanley Center for Psychiatric Research, Broad Institute of MIT and Harvard, Cambridge, MA, USA · Department of Systems and Computational Biomedicine and WorldQuant Initiative, Weill Cornell Medicine, New York, NY, USA · San Diego Supercomputer Center, University of California San Diego, La Jolla, CA, USA · Weill Institute for Neurosciences, Department of Neurology, University of California San Francisco, USA · Science for Life Laboratory, Department of Gene Technology, KTH Royal Institute of Technology, Stockholm, Sweden · European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Hinxton CB10 1SD, UK

Abstract

While AI holds the potential to revolutionize space life sciences, realizing this promise is contingent upon the systematic restructuring of heterogeneous spaceflight biological data into machine-actionable, AI-ready forms. Even though open access principles support human reuse and scientific reproducibility, this does not necessarily enable AI systems to access and analyze such a diverse set of scientific datasets. In addition, the growing array of AI approaches places distinct demands on data structure, metadata, and access interfaces. In order to respond to such growing changes we propose a three-tier approach, proceeding from FAIR to AI-ready to space-ready data. We discuss existing infrastructures and how they can be improved to close the AI access gap. We conclude by proposing a neutral international coordinating body as the governance backbone for the trustworthy, agent-accessible space biology infrastructure that deep space biological research will require.

Explore similar work

Sep 15, 2026cs.RO

Artificial Intelligence-Enabled Space Robot Operations: Technologies, Challenges and Prospects

Space robots are increasingly expected to perform long-duration, contact-rich, and multi-stage operations with limited human intervention. Recent advances in artificial intelligence (AI), robot learning, and embodied foundation models provide new opportunities to improve the autonomy and adaptability of such systems, but their transfer to space is constrained by scarce mission data, space-specific dynamics and sensing conditions, limited onboard resources, and stringent safety requirements. This article reviews artificial intelligence-enabled space robot operations (AI-SRO) from a capability-building perspective. We first summarize representative operational scenarios, autonomy trends, and space-specific constraints. We then establish a three-layer technical framework comprising capability foundations, capability formation, and capability deployment/evolution. Within this framework, we review simulation environments, datasets and benchmarks; task and environment understanding, state perception, decision-making and planning, and action execution; and onboard deployment, ground-to-space adaptation, continual learning, and capability transfer. Finally, we propose key research directions toward trustworthy simulation and data, open-world multimodal cognition, long-horizon safe decision-making, physically constrained policy learning, and space computing infrastructures.
Zeyuan Huang, Gang Chen, Zixuan Hao +7
Jul 2, 2026cs.AI

Automated Data Readiness for Scientific AI

Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenance-instrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.
Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi +8
Apr 29, 2026cs.AI

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

AI-for-Science (AI4Science) is increasingly transforming scientific discovery by embedding machine learning models into prediction, simulation, and hypothesis generation workflows across domains. However, the effectiveness of these models is fundamentally constrained by the AI-readiness of scientific data, for which no scalable and systematic evaluation mechanism currently exists. In this work, we propose SciHorizon-DataEVA, a novel agentic system to scalable AI-readiness evaluation of heterogeneous scientific data. At the evaluation-criteria level, we introduce the Sci-TQA2 principles, which organize AI-readiness into four complementary dimensions: Governance Trustworthiness, Data Quality, AI Compatibility, and Scientific Adaptability. Each dimension is decomposed into measurable atomic elements that enable fine-grained and executable assessment. To operationalize these principles at scale, we develop Sci-TQA2-Eval, a hierarchical multi-agent evaluation approach orchestrated through a directed, cyclic workflow. Our Sci-TQA2-Eval dynamically constructs dataset-aware evaluation specifications by combining lightweight dataset profiling, applicability-aware metric activation, and knowledge-augmented planning grounded in domain constraints and dataset-paper signals. These specifications are executed through an adaptive, tool-centric evaluation mechanism with built-in verification and self-correction, enabling scalable and reliable assessment across heterogeneous scientific data. Extensive experiments on scientific datasets spanning multiple domains demonstrate the effectiveness and generality of SciHorizon-DataEVA for principled AI-readiness evaluation.
Dianyu Liu, Chuan Qin, Xi Chen +6