cs.AISep 23, 2026

BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

Authors: Eranga Bandara, Xueping Liang, Asanga Gunaratna, Tharaka Hewa, Abdul Rahman, Peter Foytik, Safdar H. Bouk, Sachini Rajapakse, +14 more

Organizations: Old Dominion University, Norfolk, VA, USA · Florida International University, USA · AI Motion Labs, Melbourne, Australia · Center for Wireless Communications, University of Oulu, Finland · Deloitte & Touche LLP, USA · Nanyang Technological University, Singapore · University of Colombo, Sri Lanka · Accenture Technology Labs, Arlington, VA, USA · GSI Scandinavia AB · Lithuanian University of Health Sciences

Abstract

DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive, judgment-intensive, inconsistent across operators, and frequently undocumented. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines. The framework decomposes the pipeline into six specialized AI agents, covering sample intake and quality control, alignment, variant calling, annotation, cross-stage monitoring, and reporting. Critically, BaseCamp agents do not perform sequence analysis: established tools execute alignment, calling, and annotation, while the agents select among them, configure them, interpret their output, and decide what follows. This confines language model reasoning to the judgment layer where it is reliable and preserves the reproducibility existing tooling guarantees. Agent reasoning is powered by a consortium of fine-tuned, domain-specialized large language models coordinated by a central reasoning LLM, executing locally so no sequencing data leaves the operating environment, under human-in-the-loop orchestration. Evaluation shows agent-generated configurations are concordant with expert practice, that an explicit filtering ledger renders inspectable what filtering otherwise removes without trace, and that cross-stage anomaly detection surfaces conditions execution monitoring misses. BaseCamp offers a generalizable blueprint for agentic automation of scientific data pipelines.

Figures & tables

Explore similar work

CardsList
  1. BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

    Jul 21, 2026Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang +9Bacterial Colony CountingArtificial Intelligence Agents

  2. EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

    Jun 11, 2026Harihara Muralidharan, Reema Baskar, Soo Hee Lee +2Artificial Intelligence Agents

  3. From Research Question to Scientific Workflow: Leveraging Agentic AI for Science Automation

    Apr 23, 2026Bartosz Balis, Michal Orzechowski, Piotr Kica +2Scientific WorkflowsAgentic Workflows