GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation
Authors: Tanmoy Kanti Halder, Akash Ghosh, Arijit Roy, Sriparna Saha
Organizations: Department of Computer Science and Engineering, Indian Institute of Technology Patna, Bihta, 801106, Bihar, India · Prasannadeb Women’s College, India
Large language models (LLMs) have demonstrated strong capabilities in biological reasoning; however, genomic disease inference remains largely dependent on memorized gene-disease associations rather than understanding biological pathways. This shortcut learning undermines robustness and generalization, and breaks down when molecular identifiers are unavailable. We present GenoMorph, a multimodal genomic reasoning framework that shifts disease prediction from associative gene-disease mapping toward pathway-grounded reasoning. GenoMorph couples a frozen DNA foundation model with question-conditioned cross-attention fusion, self-adaptive latent reasoning (LatentSp), a residual reasoning gate for iterative genomic evidence reinjection, and rejection sampling fine-tuning regularized by hierarchical optimal transport (OT). Rather than learning direct gene-disease mappings, GenoMorph aligns genomic sequence representations with latent pathway dynamics, enabling reasoning trajectories that follow molecular interactions before producing disease predictions. LatentSp dynamically allocates computation according to reasoning confidence, reducing unnecessary reasoning steps and improving inference efficiency. We further construct an anonymized benchmark from the Kyoto Encyclopedia of Genes and Genomes (KEGG), replacing every gene and molecular identifier with anonymous symbols while preserving sequences and pathway topology, thereby removing memorization shortcuts. GenoMorph raises the weighted F1 from 0.7863 (BioReason) to 0.9412, and rejection sampling fine-tuning with self-adaptive latent reasoning pushes it to 0.9725 while cutting latency nearly 60%. On the anonymized benchmark it reaches 0.9465 F1, substantially outperforming prior systems and confirming that accurate disease prediction can arise from pathway reasoning rather than memorized gene-disease associations.
Figures & tables
Figure 1: Training and inference pipeline of GenoMorph . The framework begins by encoding genomic DNA sequences using the frozen Evo2-7B foundation model to obtain DNA embeddings, which are combined with clinical question embeddings through question-conditioned CrossAttentionFusion (Stage 1) to generate question-aware genomic representations. Then it performs curriculum-based LatentSp training to initialize latent reasoning, followed by ThinkingResidualGate pretraining for controlled genomic evidence reinjection. Stage 2 computes offline Hierarchical Optimal Transport (HiRef-OT) target manifolds and correspondences ( dOT ), providing geometry-aware rewards for reinforcement optimization. Stage 3 first applies GRPO to optimize latent reasoning trajectories, genomic evidence reinjection, and language generation using the HiRef-based reward. The resulting reasoning traces are subsequently filtered and used for Rejection Sampling Fine-Tuning (RSFT), enabling the model to internalize the adaptive reasoning policy without requiring explicit per-step entropy estimation or a learned entropy threshold during inference. The final self-adaptive latent model dynamically allocates computation by continuing latent reasoning for difficult samples while reinjecting genomic evidence only when necessary, producing open-ended disease predictions grounded in biologically coherent pathway reasoning rather than direct gene–disease associations.
Component
Setting
Backbone LLM
Qwen3-1.7B-Instruct
Genomic encoder
Evo2-7B (frozen)
DNA embedding layer
blocks.28.mlp.l3
Fusion module
CrossAttentionFusion
Trainable parameters
CrossAttentionFusion + LoRA adapters
LoRA rank / α / dropout
32 / 64 / 0.05
Table 1: Training configuration for supervised learning (Stage 1).
Parameter
Value
Initialization
Best supervised checkpoint
Optimizer
AdamW
Learning rate
2×10−6
Batch size
1 (effective = 4)
LoRA ( r , α )
16, 32
Generations
8
Table 2: Training configuration for Group Relative Policy Optimization (GRPO) (Stage 3).
Table 5: Overall performance on the genomic disease reasoning benchmark. Results are reported on both the original benchmark (Named) and the anonymized benchmark (Anonymous). Macro-F1 (M), Weighted-F1 (W), and average inference time per example are reported. Best results are shown in bold .
Figure 2: Accuracy comparison on the named and anonymized genomic disease reasoning benchmarks. The vertical distance between the paired markers represents the performance degradation caused by removing explicit gene and molecular entity names. GenoMorph exhibits the smallest performance gap, demonstrating robust sequence-level genomic reasoning under entity anonymization.
Figure 3: Average inference time per sample on the named and anonymized genomic disease reasoning benchmarks. Lower inference time indicates higher computational efficiency. GenoMorph with self-adaptive RSFT achieves the lowest inference latency while simultaneously providing the highest predictive performance.
Model
Δ Accuracy ↓
Accuracy Retention (%) ↑
LLM-only (Qwen3)
39.66
55.4
BioReason (SFT)
30.35
66.5
BioReason (GRPO)
31.03
62.5
GenoMorph (Ours)
3.45
96.4
GenoMorph+ (Self-Adaptive RSFT)
2.76
97.2
Table 6: Robustness to gene-name anonymization. Absolute decrease in accuracy and accuracy retention after replacing all gene and molecular entity names with globally consistent anonymous identifiers. Lower accuracy drops and higher retention indicate stronger sequence-level genomic reasoning.
Figure 4: Anonymous benchmark accuracy and accuracy retention after gene-name anonymization. Gray bars denote baseline models, whereas blue bars correspond to the proposed GenoMorph variants. The red line represents accuracy retention relative to the named benchmark. GenoMorph and GenoMorph +RSFT achieve both substantially higher anonymous accuracy and significantly greater accuracy retention, demonstrating robust sequence-level genomic reasoning under entity anonymization.
\multirow 2* Variant
\multirow 2* DNA Fusion
\multirow 2* LatentSp
\multirow 2* Latent Reasoning
\multirow 2* Gate
\multirow 2* HiRef OT
Named Benchmark
Anonymous Benchmark
Acc
Weighted-F1
Time (s)
Acc
Weighted-F1
Time (s)
Stage 1: CrossAttn SFT
CrossAttn
✗
✗
✗
✗
0.9310
0.7573
43.85
0.5586
0.4762
26.84
Stage 1.5: +LatentSp
CrossAttn
✓
✗
✗
✗
0.9552
0.9538
15.05
0.9034
0.8952
17.10
Stage 1.5: +Gate
CrossAttn
✓
✗
✓
✗
0.9828
0.9793
14.78
0.9069
0.9073
17.52
GRPO ( θlow=0 )
CrossAttn
✓
No Latent
✓
✓
0.9379
0.8690
29.24
0.7518
0.7074
30.82
GRPO ( θlow=3 )
CrossAttn
✓
Full Latent
✓
✓
0.9448
0.9067
17.11
0.7459
0.7362
18.45
Table 7: Ablation study of GenoMorph . The table shows the contribution of each architectural component, including Cross-Attention DNA fusion, LatentSp curriculum learning, latent reasoning, adaptive gating, HiRef OT reward, and self-adaptive RSFT.
Figure 5: Optimization trajectory of GenoMorph throughout the ablation study in the accuracy–efficiency plane. Each point corresponds to one training stage or architectural variant. The blue dashed curve denotes the optimization trajectory as successive components are incorporated, while the red solid line represents the Pareto frontier. The red dashed segment illustrates the performance recovery obtained by introducing the HiRef OT reward compared with its ablated counterpart. The shaded green region indicates the desirable operating regime with simultaneously high prediction accuracy and low inference time. The final self-adaptive RSFT model achieves the best trade-off between predictive performance and computational efficiency.
Figure 6: Qualitative comparison on the original benchmark. GenoMorph produces biologically grounded reasoning and correctly predicts the target disease, whereas GPT-4o-mini exhibit wrong reasoning hallucinations leading to incorrect diagnoses and BioReason hallucinates to partially correct diagnoses.
Figure 7: Qualitative comparison after gene-name anonymization. Despite the removal of explicit gene identifiers, GenoMorph correctly infers the disease by reasoning from genomic sequence information, whereas GPT-4o-mini and BioReason fail under the anonymized setting.
Disease
Misses
Observation
Creutzfeldt–Jakob disease
7
Predicted broader category (Prion disease)
Sphingolipidoses
1
Predicted specific subtype
Hepatocellular carcinoma
2
Genuine prediction errors
Pancreatic ductal adenocarcinoma
1
Genuine prediction error
Colorectal cancer
1
Genuine prediction error
Non-small cell lung cancer
1
Genuine prediction error
Table 8: Manual analysis of the remaining prediction errors produced by GenoMorph with self-adaptive RSFT.
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Ada Fang, Nikitha Thoduguli, Lukas Fesser +3
Harvard University · Massachusetts Institute of Technology
Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p < 0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.
Dominic Okonkwo, Adetayo Okunoye, Ismailcem Budak Arpinar
Multi-omics sequences contain complex biological patterns, yet deciphering their mechanisms for automated scientific discovery remains challenging. As large language models (LLMs) interpret these sequences, evaluating both predictions and scientific reasoning is critical. However, existing benchmarks for multi-omics sequence tasks rely on classification and regression metrics, neglecting whether models grasp the underlying biological evidence. We introduce OmicsBench, the first reasoning benchmark for multi-omics sequences, comprising 1,160 expert-validated questions across six tasks spanning DNA regulation, RNA processing, and protein function. OmicsBench requires traceable evidence chains, evaluated using instance-specific rubrics developed with domain experts. Evaluating 17 LLMs reveals an inverse relationship: while scientific LLMs outperform general-purpose LLMs in sequence classification accuracy, they fail to provide valid evidence to support their predictions. One plausible interpretation is shortcut learning: specialized models may rely on statistical patterns rather than the biological mechanisms needed for scientific discovery. Motivated by this finding, we introduce tool-augmented on-policy distillation (TA-OPD), a post-training method to align sequence prediction with evidence-grounded biological reasoning. Across five Qwen3.5 models spanning 0.8B to 27B parameters, TA-OPD consistently strengthens biological evidence grounding while improving predictive performance on most tasks. These gains persist across model scales, indicating that stronger sequence reasoning does not arise solely from increased model capacity, but can be improved through evidence-aware training. Together, OmicsBench and TA-OPD provide a framework for diagnosing reasoning failures in multi-omics LLMs and a path toward models whose predictions are better grounded in biologically meaningful evidence.
Jie Ying, Zhefan Wang, Zihong Chen +11
Shanghai Artificial Intelligence Laboratory · Yazhouwan National Laboratory · Shanghai Innovation Institute +2