Authors: Jens Lundsgaard, Colin Mikulski, Zhixuan Yan, Dhananjay Bhaskar
Organizations: Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI, USA · Department of Mathematics, University of Wisconsin–Madison, Madison, WI, USA · Department of Biomedical Engineering, University of Wisconsin–Madison, Madison, WI, USA · Biophysics Graduate Program, University of Wisconsin–Madison, Madison, WI, USA · Data Science Institute, University of Wisconsin–Madison, Madison, WI, USA · Center for Genomic Science Innovation, University of Wisconsin–Madison, Madison, WI, USA · Wisconsin Institute for Translational Neuroengineering, University of Wisconsin–Madison, Madison, WI, USA
Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.
Figures & tables
Figure 1: SPINET architecture. (a) Backbone coordinates from a molecular dynamics trajectory are converted into residue-level node features and geometric edge features. (b) Sheaf message passing encodes residue interactions within each frame, an RNN aggregates these representations over time, and additional sheaf message-passing blocks refine them before amino acid classification.
Figure 2: SPINET learns a restriction map Fj≤eij for each incident node–edge pair. During sheaf message passing, the map transforms the neighbor representation hj into the edge space as Fj≤eijhj , where it is combined with the learned edge representation gij . The resulting message is mapped back to the node space and used to update the residue representation hi ; the updated representation is then concatenated with the previous hi .
Paradigm
Model
Params
top-1 ↑
top-5 ↑
top-10 ↑
Perplexity ↓
Dynamic
SPINET
1.5M
0.567 ± 0.060
0.898 ± 0.038
0.979 ± 0.017
3.736 ± 0.676
TGNN
1.4M
0.319 ± 0.054
0.712 ± 0.058
0.898 ± 0.039
8.872 ± 1.592
Traj–MLP (full)
23.1K
0.406 ± 0.063
0.774 ± 0.053
0.935 ± 0.031
6.301 ± 1.001
Traj–MLP (dihedrals)
22.5K
0.286 ± 0.059
0.643 ± 0.062
0.860 ± 0.046
9.848 ± 1.438
Traj–MLP (coords)
22.7K
0.369 ± 0.060
0.724 ± 0.052
0.904 ± 0.036
7.297 ± 1.063
Ensemble
DynamicMPNN (re-trained)
4.2M
0.376 ± 0.063
0.754 ± 0.054
0.919 ± 0.034
7.242 ± 1.228
Table 1: Native-sequence recovery on the 320K mdCATH test split. Top- k recovery is the fraction of residues whose native amino acid appears among the k highest-ranked predictions; perplexity measures the probability assigned to native residues. Higher recovery and lower perplexity indicate better performance. Parameter counts are shown for each model.
Sequence source
TM-score ↑
lDDT ↑
RMSD (Å) ↓
TM tar/dec↑
RMSD tar/dec↓
pLDDT tar/dec↑
Target favored (%) ↑
Native
0.824
0.835
2.405
7.244
0.153
1.44
99.0
SPINET
0.760
0.777
2.976
6.912
0.194
1.36
99.5
SPINET (Bo5) †
0.810
0.804
2.360
7.370
0.154
1.42
99.6
TGNN
0.619
0.693
4.576
6.041
0.286
1.16
99.1
Traj–MLP (full)
0.209
0.438
13.231
2.324
0.772
1.03
91.1
MapDiff
0.729
0.765
3.323
6.473
0.219
1.18
97.7
Table 2: Structural recovery and target specificity on the 320K mdCATH test set. TM-score, lDDT, and C α RMSD compare AF3 predictions with representative target conformations. Target/decoy ratios compare recovery under target and unrelated decoy templates; “target favored” is the percentage of domains with a higher TM-score under the target template. Values are medians across the 788 domains evaluable for every method. Higher TM-score, lDDT, and target/decoy TM-score and pLDDT ratios, together with lower RMSD and target/decoy RMSD ratio, indicate better performance. Interquartile ranges and additional metrics appear in Appendix B .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Sequence source
TM-like ↑
lDDT ↑
RMSD (Å) ↓
pLDDT ↑
All states
Native
0.824 (0.615–0.922)
0.835 (0.751–0.887)
2.405 (1.486–4.529)
83.30 (77.08–87.71)
SPINET
0.760 (0.540–0.871)
0.777 (0.688–0.818)
2.976 (2.054–5.457)
73.17 (69.17–76.82)
SPINET (Bo5)
0.810 (0.643–0.901)
0.804 (0.735–0.841)
2.360 (1.736–3.892)
73.97 (69.41–77.91)
TGNN
0.619 (0.362–0.779)
0.693 (0.589–0.744)
4.576 (3.023–8.059)
67.93 (63.78–73.05)
Traj–MLP (full)
0.209 (0.105–0.478)
0.438 (0.358–0.578)
13.231 (8.425–17.190)
61.44 (55.00–66.85)
MapDiff
0.729 (0.457–0.861)
0.765 (0.667–0.820)
3.323 (2.170–5.759)
76.81 (71.34–81.48)
Appendix
Table 3: AF3 template-based structural recovery on the 320K mdCATH test set. The all-states analysis reports domain-level median (IQR) over the 788 evaluable domains for every method. Sensitivity analyses retain representative conformations for which the native sequence achieves a TM-score above 0.5 or 0.7 . Best-of-five (Bo5) selects the SPINET candidate with the highest mean TM-score within each subset. Higher TM-score, lDDT, and pLDDT and lower RMSD indicate better recovery.
Sequence source
Decoy TM-like ↑
Decoy lDDT ↑
Decoy RMSD (Å) ↓
Decoy pLDDT ↑
n domains
Native
0.104 (0.085–0.125)
0.302 (0.245–0.364)
16.157 (13.565–19.030)
55.04 (42.16–72.48)
796
SPINET
0.106 (0.089–0.122)
0.301 (0.245–0.364)
15.990 (13.547–18.838)
52.22 (41.08–63.39)
796
TGNN
0.102 (0.084–0.118)
0.297 (0.244–0.356)
16.731 (14.031–19.558)
56.61 (44.12–67.05)
796
Traj–MLP (full)
0.091 (0.072–0.109)
0.274 (0.233–0.333)
18.621 (16.117–21.673)
53.90 (43.87–63.31)
796
MapDiff
0.105 (0.084–0.126)
0.303 (0.248–0.362)
16.029 (13.524–19.448)
62.18 (47.76–72.84)
788
PiFold
0.105 (0.086–0.125)
0.306 (0.248–0.365)
16.056 (13.471–19.297)
61.70 (46.78–73.64)
788
Appendix
Table 4: AF3 structural recovery using templates unrelated to the target conformations. Values are domain-level median (IQR); the number of evaluable domains is shown for each method. Low TM-scores and high RMSDs indicate weak recovery of the decoy structures.
Sequence source
Target/decoy ratio
Favored (%)
n domains
TM ↑
lDDT ↑
RMSD ↓
pLDDT ↑
TM
lDDT
RMSD
pLDDT
Native
7.244 (5.586–8.900)
2.654 (2.035–3.475)
0.153 (0.085–0.313)
1.44 (1.05–1.94)
99.0
98.5
98.5
87.1
788
SPINET
6.912 (5.275–8.271)
2.439 (1.875–3.206)
0.194 (0.119–0.370)
1.36 (1.13–1.78)
99.5
98.6
98.9
95.7
788
SPINET (Bo5)
7.370 (5.706–8.865)
2.595 (1.964–3.326)
0.154 (0.097–0.282)
1.42 (1.14–1.83)
99.6
98.9
99.6
94.5
788
TGNN
6.041 (3.836–7.391)
2.201 (1.670–2.871)
0.286 (0.175–0.536)
1.16 (1.04–1.49)
99.1
97.0
98.2
92.4
788
Traj–MLP (full)
2.324 (1.329–4.969)
1.471 (1.154–2.274)
0.772 (0.455–0.921)
1.03 (1.00–1.28)
91.1
87.2
86.4
78.7
788
Appendix
Table 5: Target-decoy structural specificity on the 320K mdCATH test set. Values are domain-level median (IQR) over the 788 domains evaluable for every method. Ratios compare recovery under target and decoy templates; “favored” gives the percentage of domains with better recovery under the target template. Higher TM-like, lDDT, and pLDDT ratios and lower RMSD ratios indicate greater target specificity.
Temperature
top-1 ↑
top-5 ↑
top-10 ↑
Perplexity ↓
450K
0.700 ± 0.082
0.958 ± 0.037
0.993 ± 0.011
2.356 ± 0.643
413K
0.644±0.074
0.934±0.039
0.988±0.014
2.841±0.663
379K
0.598±0.068
0.910±0.039
0.982±0.016
3.330±0.672
348K
0.572±0.062
0.896±0.040
0.979±0.017
3.613±0.663
320K
0.564±0.063
0.890±0.042
0.976±0.019
3.712±0.671
Appendix
Table 6: Sequence recovery and perplexity across mdCATH simulation temperatures. Higher top- k recovery and lower perplexity indicate better sequence prediction.
Paradigm
Model
Params
top-1 ↑
top-5 ↑
top-10 ↑
Perplexity ↓
Dynamic
SPINET
1.5M
0.518 ± 0.051
0.878 ± 0.033
0.972 ± 0.015
4.261 ± 0.667
TGNN
1.4M
0.231 ± 0.069
0.591 ± 0.080
0.820 ± 0.053
18.788 ± 6.644
Traj–MLP (full)
23.1K
0.310 ± 0.053
0.693 ± 0.049
0.887 ± 0.037
8.575 ± 1.126
Traj–MLP (dihedrals)
22.5K
0.268 ± 0.050
0.621 ± 0.052
0.844 ± 0.039
10.206 ± 1.305
Traj–MLP (coords)
22.7K
0.276 ± 0.044
0.672 ± 0.053
0.868 ± 0.038
9.751 ± 1.186
Ensemble
DynamicMPNN (re-trained)
4.2M
0.338±0.058
0.726±0.059
0.904±0.033
8.272±1.499
Appendix
Table 7: Sequence recovery and perplexity on the ATLAS test split for SPINET , its ablations, and static and ensemble baselines. Higher top- k recovery and lower perplexity indicate better performance.
aDepartment of Computational, Quantitative, and Synthetic Biology (CQSB), UMR 7238 IBPS, Sorbonne Universit´e, CNRS Paris, 75005, France · cUniv. Grenoble Alpes, CNRS, Grenoble INP, LJK 38000 Grenoble, France. · bInstitut universitaire de France (IUF)