cs.LGSep 30, 2026

ORBIT-FMIB: Tracking Order-Resolved Epistatic Information Through ESM-2

Authors: Maryam Rahimimovassagh, Ivan Garibay, Niloofar Yousefi

Organizations: University of Central Florida

Abstract

Protein foundation models support mutation-effect and structural prediction, but predictive performance alone does not reveal which forms of biological interaction information remain accessible through model depth. We ask whether ESM-2 retains higher-order epistatic information as strongly as first- and second-order information across its representation hierarchy, introducing ORBIT-FMIB, a diagnostic framework combining Walsh-based interaction decomposition with subset-conditioned neural dependence estimation. The method is validated on synthetic landscapes with known interaction structure before being applied to the dense four-site GB1 fitness landscape using frozen ESM-2 representations. An initial production run suggested ESM-2 retains higher-order epistatic information less well than lower-order information (ΔHO−LO=−0.107Δ_{\mathrm{HO-LO}}=-0.107). An independent replication of the complete measurement grid, under matched GPU hardware and identical critic seeds, substantially reduced this contrast (ΔHO−LO=−0.017Δ_{\mathrm{HO-LO}}=-0.017), and its sign was unstable across otherwise-defensible evaluation-pairing choices applied to the same trained critics (−0.011-0.011 to +0.015+0.015). We therefore do not currently have robust evidence that ESM-2 selectively loses higher-order epistatic information, nor that retention is equal across orders; the directional question remains open. The measurement protocol itself, including its documented removal of a positional-subset shortcut in pooled critics, remains validated and is unaffected by this finding. ORBIT-FMIB is offered as a diagnostic framework for probing interaction structure in protein foundation models; this study's own replication result illustrates why such probing requires adequately-powered reproducibility checks before its output is treated as a biological finding.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 9, 2026cs.LG

Structural Interpretations of Protein Language Model Representations via Differentiable Graph Partitioning

Protein language models such as ESM-2 learn rich residue representations that achieve strong performance on protein function prediction, but their features remain difficult to interpret as structural &\& evolutionary signals are encoded in dense latent spaces. We propose a plug-&\&-play framework that projects ESM-2 representations onto protein contact graphs &\& applies SoftBlobGIN\textbf{SoftBlobGIN}, a lightweight Graph Isomorphism Network with differentiable Gumbel-softmax substructure pooling, to perform structure-aware message passing &\& learn coarse functional substructures for downstream prediction tasks. Across enzyme classification, SoftBlobGIN achieves 92.8% accuracy &\& 0.898 macro-F1. Unlike post hoc analysis of protein language models alone, our method produces directly auditable structural explanations: GNNExplainer recovers biologically meaningful active-site residues, spatially localized functional clusters, &\& catalytic contact patterns. On binding-site detection, SoftBlobGIN improves residue AUROC from 0.8850.885 using an ESM-2 linear probe to 0.9830.983, indicating that these structural explanations are not recoverable from language-model features alone. Learned blob partitions provide an additional layer of interpretability by automatically grouping residues into functional substructures, with blobs containing annotated active-site residues showing 1.85×1.85\times higher importance than other blobs (ρ=0.339ρ{=}0.339, p=0.009p{=}0.009), without any active-site supervision. Our framework requires no retraining of the language model, adds only ∼\sim1.1M parameters, &\& generalises across ProteinShake tasks, achieving Fmax⁡F_{\max} of 0.7330.733 on Gene Ontology prediction &\& AUROC of 0.9690.969 on binding-site detection. We position this as an interpretable structural companion to protein language models that makes their predictions more transparent &\& auditable.
Sep 28, 2026cs.AI

A General Harness for Protein Foundation Model Fitness Prediction

Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.
Sep 30, 2026q-bio.BM

StabilityArc: Decoding Protein Sequence Embeddings into Generalizable Stability Landscapes

Every protein has a unique stability landscape, but the physical consequences of mutation are governed by recurring biochemical constraints. We test whether a shared decoder, trained on measurements from diverse proteins, can interpret these constraints in an unseen target, enabling cross-protein transfer for initial experimental round prescreening. We present StabilityArc , which maps frozen ESMC-600M residue representations through a shared RoPE transformer to an Lx20 matrix of substitution effects; a symmetric, contact-aware residual aids in predicting epistasis in simultaneous substitutions. In 66 strict leave-one-protein-out evaluations covering 134,794 ProteinGym variants, StabilityArc achieves 0.7134 Spearman correlation, exceeding the strongest zero-shot baseline, ProSST-2048 (0.6526), by 0.0608. We further explore the utility of this method by providing the score as a prior for Kermut, achieving Spearman correlation of 0.8280 across three supervised split schemes, improving on Kermut's reported 0.8167.