When Do Biological Reasoning Models Use Their Biological Inputs?
Organizations: Harvard University · Massachusetts Institute of Technology
Abstract
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Figures & tables
| Model | Modality | Checkpoints | Input sources | Task | Evaluation data |
| BioReason ( Fallahpour et al., 2025 ) checkpoints from ( Fesser et al., 2026 ) | Genomic sequence | SFT, RL ; 42 post-training checkpoints | Projected reference and variant Evo2 representations ( ); chromosome ( ), gene ( ), KEGG pathway ( ), and network ( ) text | Disease prediction | 1,449 KEGG-derived queries over 708 genomes and 37 disease labels; 165 genome-dependent queries; 145 held-out queries for the auxiliary task evaluation |
| ChatNT ( de Almeida et al., 2025 ) | Genomic sequence | – | Nucleotide Transformer representations ( ) | Splice donor, splice acceptor, and TATA promoter prediction | 4,578 ChatNT test queries from three binary tasks (2,000 splice donors, 2,000 splice acceptors, 578 TATA promoters; balanced, up to 1,000 per class) |
| BioReason-Pro ( Fallahpour et al., 2026 ) | Protein sequence | SFT, RL | ESM3 ( ); GO-GPT ( ); InterPro ( ); organism text ( ) | Protein function prediction | 14,102 reviewed human proteins; 91-protein temporal holdout; 540 proteins across 140 InterPro families; 96 GO NOT pairs |
| Prot2Text-V2 ( Fei et al., 2025 ) | Protein sequence | – | ESM2 ( ); protein name ( ); taxon text ( ) | Protein function description generation | 3,917 proteins from test split; 632 same-taxon pairs (1,264 evidence conflicts between name and sequence) |
| Cell2Sentence-Scale ( Levine et al., 2024 ) | Single-cell transcriptome | 2B, 27B | Expression-ranked gene sentence ( ) | Cell type annotation and rationale generation | Five atlases; 4,846 cells for annotation; 2,000 cells for rationale analysis |
| CellWhisperer ( Schaefer et al., 2025 ) | Single-cell transcriptome | – | CellWhisperer transcriptome representation ( ); expression-ranked top- gene list prefixes ( ), when supplied | Cell type prediction and evidence conflict | Same 4,846 C2S-Scale cells (five atlases); 330 same-tissue pairs, with correct prediction from representation only (660 conflicts) |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Atlas | Accession | Cells | Types | Options | Label source |
| immune1 | CELLxGENE 62ef75e4 ; E-MTAB-11536 | 30,423 | 35 | 34 | Cell Ontology names |
| immune2 | CELLxGENE e9360edf ; GSE271896 | 33,923 | 38 | 38 | Cell Ontology names |
| immune3 | CELLxGENE cc431242 ; GSE299043 | 30,593 | 31 | 31 | Cell Ontology names |
| pancreas | GSE84133 (Baron et al. 2016) | 5,641 | 14 | 14 | study shorthand |
| lung | GSE280502 (Kim et al. 2024) | 6,160 | 7 | 7 | study shorthand |
| Label transition, intact shuffled | RL | SFT |
| correct correct | (82.3) | (75.5) |
| correct incorrect | (2.4) | (6.8) |
| incorrect correct | (2.5) | (5.9) |
| incorrect incorrect, same label | (10.8) | (8.7) |
| incorrect incorrect, other label | (2.1) | (3.0) |
| Accuracy, intact | 0.847 | 0.823 |
| Quantity | Value |
| Proteins whose emitted GO set changes | (46.8%) |
| Proteins whose improves | (11.2%) |
| Proteins whose worsens | (10.8%) |
| Proteins whose is unchanged | (78.0%) |
| Mean Jaccard of the emitted GO term sets |
| Model | Condition (figure label) | Fields | |||
| BioReason | Chromosome | Network | Gene list | Gene in query | |
| all text ( wt ) | ✓ | ✓ | ✓ | ✓ | |
| pathway fields, symbols masked ( no_gene ) | ✓ | M | M (descriptions kept) | M | |
| query naming the gene ( no_pathway ) | ✓ | ||||
| query, gene symbol masked ( no_textkey ) | M | ||||
| BioReason-Pro | GO-GPT | InterPro | Organism | ||
| Query | Genes | Ribo. (%) | DEG prec. (%) | Prec.@5 (%) | Acc. |
| original query | 66.5 | 77.9 | 11.2 (16.2) | 23.2 (16.1) | 0.40 |
| Request for marker genes | 62.5 | 72.6 | 12.0 (16.4) | 22.3 (16.3) | 0.41 |
| Request for differentially expressed genes | 48.4 | 71.3 | 13.2 (16.3) | 20.6 (16.4) | 0.42 |
| Request for marker genes with exclusions | 61.6 | 69.9 | 13.3 (16.3) | 21.8 (16.3) | 0.42 |
| Same as above, instruction before the sentence | 65.8 | 74.9 | 11.6 (16.3) | 23.9 (16.1) | 0.40 |
| Annotated cell type given, request for marker genes only | 65.5 | 69.8 | 13.3 (16.2) | 20.9 (16.1) | – |
| Annotated type | Predicted type | ||||
| Cells | Prec.@5 | Prec.@all | Prec.@5 | Prec.@all | |
| Original prompt | |||||
| All | 2,000 | 21.6 (16.1) | 10.6 (16.2) | – | – |
| Correct prediction | 748 | 35.2 (19.9) | 16.9 (20.0) | 35.4 (20.6) | 17.4 (20.7) |
| Incorrect prediction | 1,252 | 13.4 (13.8) | 0 6.9 (13.9) | – | – |
| predicted type in atlas | 452 | 13.7 ( 0 9.4) | 0 5.0 ( 0 9.5) | 24.2 ( 0 9.9) | 0 8.1 (10.1) |
| Changed (%) | Accuracy | Cells | ||||
| Atlas | Label | Type | w/o rationale | w/ rationale | Correct to incorrect | Incorrect to correct |
| immune1 | 21.5 | 9.8 | 0.71 | 0.61 | 49 | 7 |
| immune2 | 62.5 | 44.0 | 0.12 | 0.08 | 27 | 8 |
| immune3 | 16.0 | 9.8 | 0.53 | 0.44 | 37 | 3 |
| pancreas | 18.0 | 9.8 | 0.65 | 0.60 | 22 | 2 |
| lung | 34.5 | 18.0 | 0.26 | 0.34 | 5 | 37 |
| Held-out property | Distinct values | Also in training | Queries with the value in training |
| Prompt text | 12 | 0 | 0 / 145 |
| Network definition | 9 | 0 | 0 / 145 |
| Gene named in the query | 8 | 7 | 139 / 145 |
| Genome (both windows) | 100 | 90 | 135 / 145 |
| Disease label | 8 | 8 | 145 / 145 |
| Disease | Edited base pair | Conflict | |||
| Training target, window | Intact | Intact | Donor’s pair | ||
| No auxiliary, 2,048 bp | 0.901 (0.026) | 0.002 (0.022) | 0.080 (0.002) | 0.005 (0.002) | 0.071 (0.005) |
| Auxiliary, 257 bp | 0.970 (0.009) | 0.021 (0.007) | 0.671 (0.041) | 0.549 (0.032) | 0.467 (0.072) |
| Auxiliary, 2,048 bp | 0.984 (0.005) | 0.028 (0.007) | 0.531 (0.115) | 0.432 (0.123) | 0.361 (0.139) |
| Text-only reference | 0.917 | 0.103 | – | ||
| Training target, window | Unseen genomes | Seen genomes |
| No auxiliary, 2,048 bp | 0.042 | 0.083 |
| Auxiliary, 257 bp | 0.148 | 0.691 |
| Auxiliary, 2,048 bp | 0.106 | 0.606 |
| Most frequent training pair (T C) | 0.111 | 0.101 |