cs.LGOct 8, 2026

RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts

Authors: Frederikke Isa Marin, Panagiotis Antoniadis, Dionysia Danai Brilli, Andreas Bjerregaard, Rachael DeVries, Yan Li, Ole Winther, Wouter Boomsma

Organizations: University of Copenhagen · Novo Nordisk A/S · Technical University of Denmark

Abstract

The genome holds the blueprint that governs the biological properties of the cell. Consequently, advancing our knowledge of genomic function is crucial both for a broader understanding of biology and for continued biomedical advances. The success of large language models on natural language and protein sequences has motivated similar efforts on genomic data. However, standard genomic language models (gLMs) often require extremely large computational resources and still fall behind traditional methods on some downstream tasks. Recently, MSA-based pretraining has been proposed as an efficient alternative, but existing models are limited to short input contexts, restricting their use to short-range tasks, such as variant effect prediction. In this work, we present RAGenome, the first retrieval-based gLM that scales pretraining to longer contexts (100×\times longer than existing MSA-based gLMs), allowing it to capture both across-species evolutionary relationships and within-species longer-range interactions. Trained on whole-genome alignments from 100 vertebrates, RAGenome substantially improves the long-range capabilities of MSA-based gLMs, raising gene finding performance from 0.45 to 0.60, while remaining competitive on purely evolutionary-based tasks like prioritizing pathogenic variants. RAGenome provides competitive gLM performance at a fraction of the training cost, unifying evolutionary modeling and long-range capabilities within a single, flexible, scalable framework. Code is available at https://github.com/PanosAntoniadis/RAGenome.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Wisteria: A Unified Multi-Scale Feature Learning Framework for DNA Language Model

    May 7, 2026Weihua Wang, Haoji Li, Feilong Bao +2Representation LearningLanguage Modeling

  2. DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks

    Jun 29, 2026Romain Karpinsky, Julien Mozziconacci, Mickaël DelceyLanguage Model PretrainingFine-Tuning

  3. GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation

    Sep 28, 2026Tanmoy Kanti Halder, Akash Ghosh, Arijit Roy +1LLM Reasoning with GraphsGenomics