Authors: Jens Lundsgaard, Colin Mikulski, Zhixuan Yan, Dhananjay Bhaskar
Organizations: Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI, USA · Department of Mathematics, University of Wisconsin–Madison, Madison, WI, USA · Department of Biomedical Engineering, University of Wisconsin–Madison, Madison, WI, USA · Biophysics Graduate Program, University of Wisconsin–Madison, Madison, WI, USA · Data Science Institute, University of Wisconsin–Madison, Madison, WI, USA · Center for Genomic Science Innovation, University of Wisconsin–Madison, Madison, WI, USA · Wisconsin Institute for Translational Neuroengineering, University of Wisconsin–Madison, Madison, WI, USA
Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.
Figures & tables
Figure 1: SPINET architecture. (a) Backbone coordinates from a molecular dynamics trajectory are converted into residue-level node features and geometric edge features. (b) Sheaf message passing encodes residue interactions within each frame, an RNN aggregates these representations over time, and additional sheaf message-passing blocks refine them before amino acid classification.
Figure 2: SPINET learns a restriction map Fj≤eij for each incident node–edge pair. During sheaf message passing, the map transforms the neighbor representation hj into the edge space as Fj≤eijhj , where it is combined with the learned edge representation gij . The resulting message is mapped back to the node space and used to update the residue representation hi ; the updated representation is then concatenated with the previous hi .
Paradigm
Model
Params
top-1 ↑
top-5 ↑
top-10 ↑
Perplexity ↓
Dynamic
SPINET
1.5M
0.567 ± 0.060
0.898 ± 0.038
0.979 ± 0.017
3.736 ± 0.676
TGNN
1.4M
0.319 ± 0.054
0.712 ± 0.058
0.898 ± 0.039
8.872 ± 1.592
Traj–MLP (full)
23.1K
0.406 ± 0.063
0.774 ± 0.053
0.935 ± 0.031
6.301 ± 1.001
Traj–MLP (dihedrals)
22.5K
0.286 ± 0.059
0.643 ± 0.062
0.860 ± 0.046
9.848 ± 1.438
Traj–MLP (coords)
22.7K
0.369 ± 0.060
0.724 ± 0.052
0.904 ± 0.036
7.297 ± 1.063
Ensemble
DynamicMPNN (re-trained)
4.2M
0.376 ± 0.063
0.754 ± 0.054
0.919 ± 0.034
7.242 ± 1.228
Table 1: Native-sequence recovery on the 320K mdCATH test split. Top- k recovery is the fraction of residues whose native amino acid appears among the k highest-ranked predictions; perplexity measures the probability assigned to native residues. Higher recovery and lower perplexity indicate better performance. Parameter counts are shown for each model.
Sequence source
TM-score ↑
lDDT ↑
RMSD (Å) ↓
TM tar/dec↑
RMSD tar/dec↓
pLDDT tar/dec↑
Target favored (%) ↑
Native
0.824
0.835
2.405
7.244
0.153
1.44
99.0
SPINET
0.760
0.777
2.976
6.912
0.194
1.36
99.5
SPINET (Bo5) †
0.810
0.804
2.360
7.370
0.154
1.42
99.6
TGNN
0.619
0.693
4.576
6.041
0.286
1.16
99.1
Traj–MLP (full)
0.209
0.438
13.231
2.324
0.772
1.03
91.1
MapDiff
0.729
0.765
3.323
6.473
0.219
1.18
97.7
Table 2: Structural recovery and target specificity on the 320K mdCATH test set. TM-score, lDDT, and C α RMSD compare AF3 predictions with representative target conformations. Target/decoy ratios compare recovery under target and unrelated decoy templates; “target favored” is the percentage of domains with a higher TM-score under the target template. Values are medians across the 788 domains evaluable for every method. Higher TM-score, lDDT, and target/decoy TM-score and pLDDT ratios, together with lower RMSD and target/decoy RMSD ratio, indicate better performance. Interquartile ranges and additional metrics appear in Appendix B .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Sequence source
TM-like ↑
lDDT ↑
RMSD (Å) ↓
pLDDT ↑
All states
Native
0.824 (0.615–0.922)
0.835 (0.751–0.887)
2.405 (1.486–4.529)
83.30 (77.08–87.71)
SPINET
0.760 (0.540–0.871)
0.777 (0.688–0.818)
2.976 (2.054–5.457)
73.17 (69.17–76.82)
SPINET (Bo5)
0.810 (0.643–0.901)
0.804 (0.735–0.841)
2.360 (1.736–3.892)
73.97 (69.41–77.91)
TGNN
0.619 (0.362–0.779)
0.693 (0.589–0.744)
4.576 (3.023–8.059)
67.93 (63.78–73.05)
Traj–MLP (full)
0.209 (0.105–0.478)
0.438 (0.358–0.578)
13.231 (8.425–17.190)
61.44 (55.00–66.85)
MapDiff
0.729 (0.457–0.861)
0.765 (0.667–0.820)
3.323 (2.170–5.759)
76.81 (71.34–81.48)
Appendix
Table 3: AF3 template-based structural recovery on the 320K mdCATH test set. The all-states analysis reports domain-level median (IQR) over the 788 evaluable domains for every method. Sensitivity analyses retain representative conformations for which the native sequence achieves a TM-score above 0.5 or 0.7 . Best-of-five (Bo5) selects the SPINET candidate with the highest mean TM-score within each subset. Higher TM-score, lDDT, and pLDDT and lower RMSD indicate better recovery.
Sequence source
Decoy TM-like ↑
Decoy lDDT ↑
Decoy RMSD (Å) ↓
Decoy pLDDT ↑
n domains
Native
0.104 (0.085–0.125)
0.302 (0.245–0.364)
16.157 (13.565–19.030)
55.04 (42.16–72.48)
796
SPINET
0.106 (0.089–0.122)
0.301 (0.245–0.364)
15.990 (13.547–18.838)
52.22 (41.08–63.39)
796
TGNN
0.102 (0.084–0.118)
0.297 (0.244–0.356)
16.731 (14.031–19.558)
56.61 (44.12–67.05)
796
Traj–MLP (full)
0.091 (0.072–0.109)
0.274 (0.233–0.333)
18.621 (16.117–21.673)
53.90 (43.87–63.31)
796
MapDiff
0.105 (0.084–0.126)
0.303 (0.248–0.362)
16.029 (13.524–19.448)
62.18 (47.76–72.84)
788
PiFold
0.105 (0.086–0.125)
0.306 (0.248–0.365)
16.056 (13.471–19.297)
61.70 (46.78–73.64)
788
Appendix
Table 4: AF3 structural recovery using templates unrelated to the target conformations. Values are domain-level median (IQR); the number of evaluable domains is shown for each method. Low TM-scores and high RMSDs indicate weak recovery of the decoy structures.
Sequence source
Target/decoy ratio
Favored (%)
n domains
TM ↑
lDDT ↑
RMSD ↓
pLDDT ↑
TM
lDDT
RMSD
pLDDT
Native
7.244 (5.586–8.900)
2.654 (2.035–3.475)
0.153 (0.085–0.313)
1.44 (1.05–1.94)
99.0
98.5
98.5
87.1
788
SPINET
6.912 (5.275–8.271)
2.439 (1.875–3.206)
0.194 (0.119–0.370)
1.36 (1.13–1.78)
99.5
98.6
98.9
95.7
788
SPINET (Bo5)
7.370 (5.706–8.865)
2.595 (1.964–3.326)
0.154 (0.097–0.282)
1.42 (1.14–1.83)
99.6
98.9
99.6
94.5
788
TGNN
6.041 (3.836–7.391)
2.201 (1.670–2.871)
0.286 (0.175–0.536)
1.16 (1.04–1.49)
99.1
97.0
98.2
92.4
788
Traj–MLP (full)
2.324 (1.329–4.969)
1.471 (1.154–2.274)
0.772 (0.455–0.921)
1.03 (1.00–1.28)
91.1
87.2
86.4
78.7
788
Appendix
Table 5: Target-decoy structural specificity on the 320K mdCATH test set. Values are domain-level median (IQR) over the 788 domains evaluable for every method. Ratios compare recovery under target and decoy templates; “favored” gives the percentage of domains with better recovery under the target template. Higher TM-like, lDDT, and pLDDT ratios and lower RMSD ratios indicate greater target specificity.
Temperature
top-1 ↑
top-5 ↑
top-10 ↑
Perplexity ↓
450K
0.700 ± 0.082
0.958 ± 0.037
0.993 ± 0.011
2.356 ± 0.643
413K
0.644±0.074
0.934±0.039
0.988±0.014
2.841±0.663
379K
0.598±0.068
0.910±0.039
0.982±0.016
3.330±0.672
348K
0.572±0.062
0.896±0.040
0.979±0.017
3.613±0.663
320K
0.564±0.063
0.890±0.042
0.976±0.019
3.712±0.671
Appendix
Table 6: Sequence recovery and perplexity across mdCATH simulation temperatures. Higher top- k recovery and lower perplexity indicate better sequence prediction.
Paradigm
Model
Params
top-1 ↑
top-5 ↑
top-10 ↑
Perplexity ↓
Dynamic
SPINET
1.5M
0.518 ± 0.051
0.878 ± 0.033
0.972 ± 0.015
4.261 ± 0.667
TGNN
1.4M
0.231 ± 0.069
0.591 ± 0.080
0.820 ± 0.053
18.788 ± 6.644
Traj–MLP (full)
23.1K
0.310 ± 0.053
0.693 ± 0.049
0.887 ± 0.037
8.575 ± 1.126
Traj–MLP (dihedrals)
22.5K
0.268 ± 0.050
0.621 ± 0.052
0.844 ± 0.039
10.206 ± 1.305
Traj–MLP (coords)
22.7K
0.276 ± 0.044
0.672 ± 0.053
0.868 ± 0.038
9.751 ± 1.186
Ensemble
DynamicMPNN (re-trained)
4.2M
0.338±0.058
0.726±0.059
0.904±0.033
8.272±1.499
Appendix
Table 7: Sequence recovery and perplexity on the ATLAS test split for SPINET , its ablations, and static and ensemble baselines. Higher top- k recovery and lower perplexity indicate better performance.
Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.
Handong Wang, Jiaxin Qi, Baisheng Lai +1
Computer Network Information Center, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Proteins move and deform to ensure their biological functions. Despite significant progress in protein structure prediction, approximating conformational ensembles at physiological conditions remains a fundamental open problem. This paper presents a novel perspective on the problem by directly targeting continuous compact representations of protein motions inferred from sparse experimental observations. We develop a task-specific loss function enforcing data symmetries, including scaling and permutation operations. Our method PETIMOT (Protein sEquence and sTructure-based Inference of MOTions) leverages transfer learning from pre-trained protein language models through an SE(3)-equivariant graph neural network. When trained and evaluated on the Protein Data Bank, PETIMOT shows superior performance in time and accuracy, capturing protein dynamics, particularly large/slow conformational changes, compared to state-of-the-art diffusion and flow-matching approaches, as well as traditional physics-based models. Our code and protocols are available at https://github.com/PhyloSofS-Team/PETIMOT.
Valentin Lombard, Julien Nguyen Van, Sergei Grudinin +1
aDepartment of Computational, Quantitative, and Synthetic Biology (CQSB), UMR 7238 IBPS, Sorbonne Universit´e, CNRS Paris, 75005, France · cUniv. Grenoble Alpes, CNRS, Grenoble INP, LJK 38000 Grenoble, France. · bInstitut universitaire de France (IUF)
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across all three models, we find a shared two-stage computational structure. In the first stage, early blocks initialize pairwise biochemical signals: features like charge propagate from sequence into pairwise representations through architecture-specific pathways. In the second stage, late blocks develop pairwise spatial features: distance and contact information accumulate in the pairwise representation. We verify these mechanisms causally by showing that steering charge and distance features induces predictable structural changes. Furthermore, these representations are functionally interchangeable: pairwise states can be linearly aligned and substituted across models. Together, these results suggest that folding trunks with different architectures, inputs, and training procedures converge on a shared representational organization for mapping sequence chemistry into spatial geometry.
Kevin Lu, Jannik Brinkmann, Stefan Huber +4
1Northeastern University · 2TU Clausthal · 3Harvard University +2