Organizations: College of Computing, Georgia Institute of Technology, Atlanta, Georgia, USA · Center for the Study of Systems Biology, Georgia Institute of Technology, Atlanta, Georgia, USA · School of Biological Sciences, Georgia Institute of Technology, Atlanta, Georgia, USA · Parker H. Petit Center for AI-Driven Health Innovation, Atlanta, Georgia, USA
AlphaFold 3 predicts protein structures with remarkable accuracy, yet how structural information emerges within the model remains poorly understood. Here, through causal interventions on internal representations and direct probing of every Pairformer block, we trace the formation of global protein geometry and identify the multiple sequence alignment (MSA) as a structural shortcut to the fold. Removing the MSA largely preserves local secondary structure while disrupting the long-range relationships that define global topology. Restoring the MSA-enriched pair representation at only forty residues recovers most of this lost organization, including at pairs never directly modified. This contribution depends on the detailed direction of the MSA module's output rather than its magnitude. The Pairformer rapidly converts this signal into global geometry: the final fold becomes recoverable by approximately block 9 of 48 for a majority of proteins, roughly twenty-seven blocks before the model's decoder can render it, whereas without the MSA it remains inaccessible for most proteins throughout the pass. Which homologs are supplied shapes this trajectory more strongly than which query is supplied; it persists for a designed query that never evolved but collapses for a shuffled sequence. Most importantly, an alignment built for a different protein that shares the fold, supplied only at the structurally corresponding columns, raises the median TM-score against experiment from 0.44 to 0.72, while the same alignment shifted a few residues along the chain performs worse than supplying no alignment at all. What AlphaFold 3 reads from an alignment is therefore a description of the fold itself, transferable between proteins that share one, rather than the query's own evolutionary history. This explains both its accuracy and the limits of what it has solved.
Figures & tables
Figure 1: Schematic of the alignment as a structural shortcut to the fold. As shown, the query sequence and its multiple sequence alignment (MSA), which provides evolutionary information from homologous proteins, are combined into the pair representation, which the Pairformer then transforms over 48 blocks. In this schematic, the difficulty of extracting the accurate fold is conveyed by the ruggedness of a terrain that the model traverses, and each point along a path marks one Pairformer layer. With a full alignment, the terrain is a smooth funnel and the global fold is already decodable from the representation after ten of the forty-eight blocks. With a subset of five distant homologs, the terrain is more rugged, and the same fold, though less accurate, is reached late in the pass by a far longer route. Without an MSA, the terrain offers no clear descent, and for many proteins the correct global fold is never reached at all. The structures shown along each path are stand-ins for the kind of prediction one would expect at that stage, in relative quality against the experimental structure, rather than readouts from any individual protein. What flattens the terrain is not the alignment’s evolutionary relationship to the query but the structure it encodes. Evolutionary information therefore does not merely improve the prediction—the structure latent within it is what makes the fold accessible, and how much of it the model receives sets how quickly the structure appears. This schematic is informative by analogy: there is no indication that AlphaFold 3 optimizes over a loss landscape, and the terrain is a proxy for the difficulty of the underlying problem.
Figure 2: Sparse restoration of the MSA recovers long-range structural organization in 8R5N. (A) Structural overlays against experiment (gray) for the full-MSA prediction, the 40-residue anchor intervention, and the no-MSA prediction (blue). Corresponding TM-scores against experiment are 0.984, 0.961, and 0.261. (B) C α distance matrices for the experimental structure, full-MSA prediction, no-MSA prediction, and anchor intervention. Long-range contact counts are 131, 136, 40, and 123, respectively. Exact β -bridge partnerships are lost without the MSA and recovered following anchor restoration.
Figure 3: The Pairformer transforms the MSA into structural geometry. (A) Projection of the evolving MSA-dependent representation onto the original MSA-module contribution (blue), alongside intermediate experimental TM-score (orange). (B) Relative contributions of the original MSA-aligned component and the orthogonal residual. (C) Generic and fold-specific contact-probe AUROC [ 29 ] across Pairformer blocks. (D) Native-contact preference with native direction preserved (blue) and its difference from the fully shuffled condition (orange). Their respective baselines are 0.5 and 0. (E) Per-protein experimental TM-scores across intermediate checkpoints in the novel and similar cohorts ( n=200 each). (F) Left: median TM-score against experiment for structures diffused independently from the saved representations at each checkpoint, with and without the MSA (see Supplementary Information Section S4). Right: cosine similarity between the pair representation at each block and the final representation of the same run. The late structural improvement is substantially greater with the MSA, while the final Pairformer block produces a pronounced representational change.
Figure 4: With the MSA, the global fold becomes decodable from the pair representation within the first ten Pairformer blocks; without it, the fold does not emerge for most proteins. (A) Linear probes trained on the similar cohort decode C α –C α distances of the final structure from the pair representation at checkpoints A and B and after each of the 48 blocks of the first Pairformer pass, for the 200 novel proteins with the full MSA (blue) or without it (orange). Panels show, from left to right and top to bottom: TM-score of the reconstructed structure against the final structure; the percentage of reconstructions that reach a TM-score of at least 0.5; long-range lDDT ( ∣i−j∣≥24 ) computed from the decoded distances; the coefficient of determination ( R2 ) of the decoded distances; the radius of gyration of the reconstruction as a ratio of its final value; and the Jaccard overlap between the long-range contact maps of the reconstruction and the final structure. Lines are medians and shaded bands are interquartile ranges. (B–C) Structures of 8R5N reconstructed from the decoded distances without the MSA (B) and with it (C) , at checkpoint B and Pairformer blocks 5, 9, 14, 24, 30, 36, 40, and 47. Each residue is colored by the distance of its C α atom from its position in the final structure after superposition: dark blue, ≤1 Å; light blue, 1–2 Å; yellow, 2–4 Å; orange, >4 Å. See Methods for probe training, reconstruction, and display refinement.
(a) Progress toward the fold
Block
Q (subsets A&B)
H
MPNN
Shuffled
Q (full MSA)
Q (no MSA)
14
0.322
0.329
0.327
0.200
0.688
0.245
18
0.333
0.343
0.343
0.193
0.670
0.249
24
0.346
0.367
0.366
0.195
0.668
0.251
30
0.450
0.461
0.499
0.191
0.756
0.262
36
0.504
0.524
0.595
0.193
0.785
0.270
Table 1: The MSA sets when and along which route the fold emerges; the query sets whether and how well. All values are median TM-scores between structures reconstructed from probe-decoded distances, for 100 novel proteins folded with two evolutionarily independent MSA subsets (A and B) of 5 distant homologs each ( ≤30% identity to the query and to the other subset). Blocks are numbered 0–47 within the first Pairformer pass. (a) Reconstructions compared against the final full-MSA prediction of the original query (Q), with the query replaced by a natural homolog (H, ∼ 60% identity), a ProteinMPNN sequence [ 32 ] design on Q’s backbone (MPNN, ∼ 45% identity), or Q’s residues shuffled, each folded with the same subsets. Full MSA and No MSA are Q’s own representations decoded as in Figure 4 . (b) Pairwise TM-scores between reconstructions of the same protein. QA vs QB changes the MSA subset while keeping the query; the remaining columns change the query while keeping the subset (A with A, B with B). The corresponding diffusion readouts are given in Supplementary Section S7.
Full MSA
No MSA
Block
Own fold
Partners
Unrelated
Own fold
Partners
Unrelated
0
0.272
0.236
0.206
0.204
0.286
0.285
9
0.555
0.320
0.145
0.243
0.239
0.219
14
0.598
0.348
0.145
0.249
0.236
0.217
30
0.696
0.407
0.138
0.249
0.235
0.207
40
0.777
0.468
0.142
0.264
0.247
0.216
Table 2: Sequence-divergent members of the same fold family converge on their shared structure only with the MSA. Probe reconstructions for 542 pairs of Sabmark proteins drawn from the same structural group (median 9.4% sequence identity, 65 aligned residues), decoded from representations generated with or without the MSA, each by the probe trained on representations of the same type. Own fold is the TM-score of each protein’s reconstruction against its own final full-MSA structure. Partners and unrelated are TM-scores between two proteins’ reconstructions on structurally aligned residues. unrelated pairs come from different groups and are aligned by position. The final row gives the agreement between the partners’ final predicted structures with and without the MSA. Values are medians.
Figure 5: An alignment built for a different protein recovers the fold, provided it is supplied with the right correspondence. (A) Two proteins sharing a fold at low sequence identity are aligned structurally (here the Sabmark domains of 1EVH, in blue, and 1MKE, in grey), and the columns of the donor’s alignment that correspond to aligned residues are written onto the target’s columns; every other column of the target is gapped. The target’s own query sequence is retained as the first row. The resulting prediction is compared against the prediction obtained with no alignment at all, attaining a TM-score of 0.93 against 0.37, so that the transplanted alignment alone rescued an otherwise degenerate fold. In the final overlay, the predicted structure of 1EVH is shown in blue and its experimental structure in white. (B) Sabmark pairs (94 targets) and (C) novel-cohort pairs (38 targets). Left: median TM-score of the probe reconstruction against the target’s own final full-MSA prediction, at each Pairformer block. Right: the percentage of targets whose reconstruction reaches a TM-score of at least 0.5. Lines are medians and shaded bands interquartile ranges; markers to the right of each panel give the diffusion readout of the block-47 representation ( CN ). Register shifted 3 displaces the transplanted columns three residues along the target chain; register scrambled sends them to a random permutation of the same columns, destroying their order as well as their placement. Both supply the identical donor alignment at the same depth and occupancy. Remaining conditions are reported in Supplementary Tables S7–S11.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Evolutionary information preferentially supports nonlocal structure, which can be recovered through sparse interventions. (A) TM-score against experiment with and without the MSA in novel and similar proteins ( n=200 each). (B) Retention of helices, strands, exact β -bridge partnerships, and contacts grouped by sequence separation, measured relative to the full-MSA prediction. (C) Contact retention by sequence separation in the novel cohort following MSA removal and restoration of the MSA-enriched pair representation at 10 or 40 contact-selected residues, or 40 randomly selected residues. (D) TM-score against experiment following interventions at four mutually disjoint sets of 40 contact-selected or conserved residues. (E) TM-score against experiment as a function of the absolute number of anchor residues—those whose MSA-enriched pair representations were restored within the no-MSA state—using contact-based, conservation-based, or random selection in both cohorts. (F) Corresponding dose response when the number of anchors is specified as a percentage of sequence length. The dashed reference lines in (E) and (F) indicate full-MSA and no-MSA prediction accuracy. (G) Advantage of contact-selected over randomly selected anchors in fractional TM-score recovery, stratified by the quality of the original no-MSA prediction. (H) Distributions of structural quality, measured by TM-score, for the no-MSA predictions, predictions following restoration at 40 contact-selected anchors, and full-MSA predictions, shown separately for the novel and similar cohorts. See Methods for cohort definitions, anchor selection, recovery calculations, and structural metrics.
Feature
Novel
Similar
Helical residues
0.84
0.85
Strand residues
0.71
0.65
Exact β -bridge partnerships
0.42
0.36
Contacts, short range ( 6≤∣i−j∣<12 )
0.45
0.43
Contacts, medium range ( 12≤∣i−j∣<24 )
0.40
0.37
Contacts, long range ( ∣i−j∣≥24 )
0.28
0.21
Appendix
Table S1: Fraction of each structural feature present in a protein’s full-MSA prediction that is retained in its no-MSA prediction. Medians across 200 proteins per cohort.
Pair category
Novel
Similar
UU (neither residue an anchor)
0.349
0.494
IU (one residue an anchor)
0.506
0.646
II (both residues anchors)
0.629
0.810
Appendix
Table S2: Fraction of the gap in balanced long-range contact-map accuracy between the no-MSA and full-MSA predictions closed by restoration at 40 contact-selected anchors, by whether the pair’s representation was directly modified. Medians across 200 proteins per cohort.
Figure S2: Cross-column evolutionary relationships modify the MSA-derived representation, whose detailed direction influences structural prediction and whose partial restoration recovers relationships beyond directly modified residue pairs. The MSA-specific module contribution is defined as M=(PBMSA−PAMSA)−(PBnoMSA−PAnoMSA) , distinct from the checkpoint-B difference ΔB used for anchor interventions. The magnitude–direction swaps instead use the raw module updates from the native and shuffled conditions. (A) Persistence of the original MSA-specific contribution across Pairformer blocks in the novel and similar cohorts, measured by its projection onto the evolving MSA-dependent pair representation. The same quantity is shown for the novel cohort in Figure 3A. (B) Cosine similarity and relative perturbation of the MSA-derived representation following within-column MSA shuffling, which preserves individual-column statistics while disrupting cross-column evolutionary relationships. (C) TM-score against the full-MSA prediction following exchange of magnitude and direction between the native and 99%-shuffled module updates. The four conditions combine native or shuffled magnitude with native or shuffled direction. (D) Per-protein direction and magnitude rescue among the 157 novel proteins meeting the minimum structural-difference criterion. Restoring the native direction recovers much of the native structural solution within the tested perturbations, whereas restoring the native magnitude alone produces little recovery. (E) Fractional recovery of balanced long-range contact-map accuracy following restoration at 40 contact-selected anchors, separated into pairs with neither residue selected (UU), exactly one residue selected (IU), or both residues selected (II). This panel accompanies Section S2.
Magnitude
Direction
TM-score vs full-MSA prediction
Native
Native
0.96
Shuffled
Native
0.96
Native
Shuffled
0.70
Shuffled
Shuffled
0.68
Appendix
Table S3: Median TM-score against the full-MSA prediction for the four recombinations of magnitude and direction between the native and 99%-shuffled MSA-module updates, novel cohort.
Figure S3: Structural features of the diffusion readouts across the first Pairformer pass. Measurements relative to the final recycled full-MSA prediction ( CN ), across Pairformer blocks 14, 18, 24, 30, 36, 40, and 44, C1 , and CN : helix formation, strand formation, exact β -bridge recovery, global structural agreement, and the radius of gyration as a ratio of its final value. The final panel overlays the median helix, strand, β -bridge, and global-agreement curves. Lines are medians and shaded bands are interquartile ranges.
Block
With MSA
Without MSA
18
0.246
0.210
40
0.650
0.320
44
0.923
0.354
Appendix
Table S4: Median TM-score against experiment for structures diffused independently from saved Pairformer representations, novel cohort. The pronounced late transition between blocks 40 and 44 occurs only when evolutionary information is available.
Feature
Median within-protein Spearman ρ
Residue burial
+0.39
Alignment gap fraction
−0.34
Conservation
+0.15
Covariation
+0.13
All features combined (cross-validated)
0.47
Appendix
Table S5: Association between per-residue features and the order in which residues become decodable. Positive values indicate earlier decoding. Burial was positively associated in every protein examined; gap fraction was negatively associated in 93%.
Block
A vs B
A vs Exp
B vs Exp
14
0.712
0.197
0.194
18
0.748
0.194
0.192
24
0.723
0.220
0.224
30
0.700
0.277
0.268
36
0.707
0.407
0.399
40
0.772
0.504
0.506
Appendix
Table S6: Median TM-scores for structures generated by the diffusion module from the representations of two evolutionarily independent MSA subsets, compared between subsets and against the experimental structure, for 100 novel proteins.
Sabmark
Novel cohort
Condition
TM
gap closed
TM
gap closed
n
Own full MSA (ceiling)
0.951
—
0.916
—
94 / 38
Own MSA, aligned columns
0.836
+0.75
0.884
+0.92
94 / 38
Partner’s MSA, aligned columns
0.717
+0.33
0.762
+0.43
94 / 38
Partner’s MSA, ungapped
0.743
+0.46
0.762
+0.43
94 / 38
Partner’s MSA, contiguous blocks
0.846
+0.79
—
—
20 / —
Appendix
Table S7: Every transplant condition and null. Median TM-score against the experimental structure, and the median fraction of the gap between the no-MSA and own-full-MSA predictions that each condition closes. Conditions run on only a subset of pairs are marked accordingly in n and are not directly comparable with the pooled values. Negative fractions indicate conditions that perform worse than supplying no alignment at all.
Block
Own full
Own aligned
Partner aligned
Partner ungapped
Contiguous
No MSA
0
0.274
0.218
0.224
0.233
0.240
0.203
5
0.433
0.282
0.278
0.299
0.324
0.227
9
0.557
0.326
0.319
0.362
0.432
0.243
14
0.601
0.355
0.342
0.426
0.506
0.248
24
0.612
0.389
0.356
0.428
0.517
0.248
30
0.698
0.494
0.410
0.474
0.558
0.251
Appendix
Table S8: Sabmark: fold decodability across the Pairformer, transplant conditions. Median TM-score of the probe reconstruction against each target’s own final full-MSA prediction, which is therefore 1.000 for that condition by construction. CN is the diffusion readout of the block-47 representation. The contiguous-block transplant was run only on the 20 higher-coverage targets.
Block
Own full
Shift 3
Shift 7
Scrambled
Masked
Corrupted
Complement
No MSA
0
0.274
0.210
0.198
0.211
0.189
0.202
0.186
0.203
5
0.433
0.224
0.210
0.223
0.197
0.210
0.208
0.227
9
0.557
0.243
0.231
0.243
0.185
0.214
0.233
0.243
14
0.601
0.232
0.219
0.236
0.195
0.220
0.219
0.248
24
0.612
0.220
0.200
0.221
0.194
0.215
0.204
0.248
30
0.698
0.234
0.220
0.233
0.249
0.243
0.218
0.251
Appendix
Table S9: Sabmark: fold decodability across the Pairformer, nulls. Each null supplies the same donor alignment at the same depth and occupancy, altering only where its columns land. Scores are against each target’s own final full-MSA prediction, as in Table S8 . No null produces a decodable fold in any protein before block 27.
Block
Own full
Own aligned
Partner aligned
Partner ungapped
No MSA
Corrupted
Shift 3
Scrambled
0
0.267
0.243
0.251
0.253
0.230
0.231
0.242
0.232
5
0.404
0.348
0.290
0.294
0.230
0.218
0.250
0.230
9
0.683
0.480
0.424
0.431
0.245
0.247
0.270
0.251
14
0.700
0.553
0.477
0.483
0.269
0.243
0.270
0.266
24
0.707
0.578
0.504
0.497
0.287
0.255
0.275
0.267
30
0.783
0.662
0.555
0.552
0.299
0.275
0.315
0.279
Appendix
Table S10: Novel cohort: fold decodability across the Pairformer. Transplant conditions and nulls for the 38 post-cutoff targets, scored against each target’s own final full-MSA prediction. CN is the diffusion readout of the block-47 representation.
Coverage
n
Median
Aligned
No MSA
Transplant
Own aligned
Own full
Gap
quartile
coverage
residues
closed
0.26–0.51
24
0.41
69
0.439
0.567
0.721
0.958
+0.05
0.51–0.60
23
0.56
87
0.389
0.690
0.808
0.945
+0.44
0.60–0.72
24
0.67
80
0.543
0.735
0.848
0.940
+0.32
0.72–0.93
23
0.79
80
0.451
0.856
0.907
0.953
+0.79
Appendix
Table S11: The transplant recovers as much of the fold as the donor’s alignment covers. Sabmark targets binned by the fraction of the target chain covered by the structural alignment. Median TM-scores against the experimental structure. Across targets, the Spearman correlation between coverage and transplant accuracy is +0.51 ( p=1×10−7 ), and between sequence identity and transplant accuracy +0.35 ( p=6×10−4 ). Controlling for identity, coverage retains a partial correlation of +0.45 ; controlling for coverage, identity retains +0.25 . Coverage is therefore the stronger determinant.
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across all three models, we find a shared two-stage computational structure. In the first stage, early blocks initialize pairwise biochemical signals: features like charge propagate from sequence into pairwise representations through architecture-specific pathways. In the second stage, late blocks develop pairwise spatial features: distance and contact information accumulate in the pairwise representation. We verify these mechanisms causally by showing that steering charge and distance features induces predictable structural changes. Furthermore, these representations are functionally interchangeable: pairwise states can be linearly aligned and substituted across models. Together, these results suggest that folding trunks with different architectures, inputs, and training procedures converge on a shared representational organization for mapping sequence chemistry into spatial geometry.
Kevin Lu, Jannik Brinkmann, Stefan Huber +4
1Northeastern University · 2TU Clausthal · 3Harvard University +2
AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. How biological information changes as it crosses this architectural boundary remains poorly understood. We analyze per-residue activations from the Pairformer trunk and diffusion module of Boltz-1 using linear probes, sparse autoencoders (SAEs), and causal interventions. From the trunk, both geometry (secondary structure, disorder) and sequence chemistry (amino-acid identity, signal peptides, disulfide-bond annotations) are linearly decodable. In the diffusion module, the two diverge. Secondary structure transfers essentially unchanged, whereas sequence chemistry is strongly attenuated. We then test whether decodable directions can steer the model, intervening on the final trunk single representation that conditions the diffusion module. Helix and coil directions change predicted structure dose-dependently against matched-norm random controls, but a beta-strand direction that is highly predictive (F1 =0.82) produces no measurable increase in strand content: linear decodability does not imply causal influence at the site we tested. The same probes also score markedly lower against sparse SwissProt annotations than against dense DSSP labels, because unannotated residues that the model gets right are charged as false positives; such scores are therefore lower bounds. Finally, supervised probes outscore single SAE features wherever a label already exists. We release the trained trunk and diffusion SAEs, Boltz-1 per-residue activations, and the analysis code.
Piotr Jedryszek, Tongmeng Xie, Adam Winnifrith +5
Department of Biology, University of Oxford, Oxford, UK · Evolvere Biosciences, London, UK · Independent Researcher +4
Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.
Yifan Feng, Guanjie Cheng, Shihui Ying +2
1{School of Software, BNRist, THUIBCS}, Tsinghua University, Beijing 100084, China · School of Software Technology, Zhejiang University, Ningbo 315100, China · 3Shanghai Institute of Applied Mathematics and Mechanics, School of Mechanics and Engineering Science, Shanghai University, Shanghai 200072, China +1