From Surfaces to Volumes: Registered Geometry for Protein Representation Learning
Authors: Siyuan Chen, Cai Zhou, Jinrui Zhang, Zhaokang Liang, Taku Komura, Wojciech Matusik, Stephen Bates, Tommi Jaakkola, +3 more
Organizations: University of British Columbia · Massachusetts Institute of Technology · Carnegie Mellon University · Northeastern University · The University of Hong Kong
Existing protein geometry models typically represent molecular surfaces using local geometric features such as sampled points, normals, and curvature. While effective for capturing exposed molecular shape, these representations do not explicitly model the volumetric organization beneath the surface or provide a consistent coordinate system for residue-wise volumetric structure. We introduce Protein-TetSphere, a registered residue-wise volumetric representation for proteins. Each protein chain is tetrahedralized to obtain local volumetric regions associated with individual residues, which are then registered to a shared fixed-topology tetrahedral reference and represented in a common Laplacian basis. This registration establishes consistent volumetric coordinates across residues, enabling local three-dimensional deformation to be integrated with surface and chemical information in a multimodal protein representation. We evaluate Protein-TetSphere on ligand-binding pocket classification, protein--protein interface prediction, and de novo protein binder design. Across the three tasks, Protein-TetSphere improves ligand-binding pocket balanced accuracy from 0.795 to 0.826, Pinder-Pair/Site AUROC from 0.914/0.852 to 0.932/0.866, and binder-design success from 14.95% to 19.90% on the BoltzGen Challenge Set and from 27.62% to 32.19% at the ProtDBench backbone level. These results show that registered volumetric geometry provides complementary spatial information beyond molecular surfaces across protein recognition, interaction, and design.
Figures & tables
Figure 1: Motivation of Protein-TetSphere. Left: SFTI-1 and BbKI illustrate how structurally distinct binders can engage the same trypsin target through similar target-facing interfaces while differing substantially in their residue-wise volumetric organization. Right: Protein-TetSphere is evaluated on ligand-binding pocket classification, protein–protein interface prediction, and de novo protein binder design.
Figure 2: TetSphere fitting and multimodal pretraining. Top: per-residue TetSpheres are fitted (steps 0 – 3000 ) and represented by Laplacian coefficients Cr . Bottom: surface, chemical, and TetSphere features are fused with nucleotide, ligand, and bond inputs, processed by a Pairformer-style single–pair trunk, and pretrained with geometry, identity, and pairwise structure objectives.
Figure 3: Truncated-basis reconstruction. Fitted TetSpheres and leading- M Laplacian reconstructions for Tyr35 of 3KH5 and Leu380 of 4ZE2. Colors indicate per-vertex error; values report mean error relative to residue extent.
Method
Balanced accuracy ↑
AtomSurf ( Mallet et al., 2025 )
0.795±0.005
+ Ours (w/o TetSphere)
0.807±0.003
+ Ours
0.826±0.004
Table 1: Test balanced accuracy for ligand-binding pocket classification, reported as the mean and standard deviation over five random seeds.
Figure 4: Shape comparison in the 6BD9 FAD pocket. FAD, ADP, and HEM are compared within the same pocket; numbers report ligand heavy atoms inside the fitted TetSphere shells.
Figure 5: Representative Pinder-Pair interface predictions from our model. Chain A is shown in red and chain B in blue, with TetSphere volumes highlighting residues involved in predicted contacts. Each panel shows 28–30 predicted contacts, all of which are true contacts under the 5 Å heavy-atom criterion, illustrating the correctness of our high-confidence contact predictions in these examples.
Method
Pinder-Pair ↑
Pinder-Site ↑
AtomSurf ( Mallet et al., 2025 )
0.914±0.002
0.852±0.002
+ Ours (w/o TetSphere)
0.919±0.002
0.855±0.001
+ Ours
0.932±0.001
0.866±0.001
Table 2: Protein–protein interface prediction on clustered PINDER in the holo setting. Results are test AUROC (mean ± s.d.) over five seeds.
+Ours
Criterion or stage
BoltzGen
+Cont.
(w/o Tet)
Ours
ProtDBench: independent pass rate
Normalized pLDDT
94.39 ± 0.15
95.18 ± 0.17
90.49 ± 0.22
93.97 ± 0.19
Interface pTM
26.18 ± 0.43
27.16 ± 0.42
31.77 ± 0.43
32.39 ± 0.47
Interface PAE
18.23 ± 0.38
19.60 ± 0.38
21.42 ± 0.38
23.35 ± 0.42
ProtDBench: cumulative survival
Table 3: Binder-design success across four training arms. Rows prefixed with + are cumulative, with a shared denominator within each block. Denominators are 38,400 ProtDBench sequences, 4,800 ProtDBench backbones, and 2,000 Challenge Set candidates. Best and second-best values are shaded dark and light green. The ± values are bootstrap standard deviations over designs with targets fixed. See Appendix E for details.
Figure 6: Target-level binder-generation success rates, with the same four arms as Table 3 : the released BoltzGen checkpoint, the continued-training control, the no-TetSphere control, and Ours. The left panel reports ProtDBench at the backbone level, where a backbone counts as a success if any of its eight redesigned sequences is accepted; the right panel reports final hard-filter success on the BoltzGen Challenge Set, where each candidate carries a single sequence. These correspond to the backbone-level success row and to the final cumulative row of the Challenge Set block in Table 3 . Both panels share a common vertical scale so rates can be read across benchmarks, and both retain all ten targets rather than collapsing the comparison to one aggregate; response is strongly target-dependent in both.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Stream
Input and encoder
Mask policy
Reconstruction target
Surface
16 local points and four descriptors; three residue-local invariant message-passing blocks
15% of valid residues
Four surface descriptors (Smooth-L1)
Chemistry / AA
Hydropathy and charge; input projection followed by two pair-biased encoder layers
Joint 20% AA–chemistry mask
Amino-acid identity (20-class cross entropy)
TetSphere
64 three-channel coefficients; mode/eigenvalue embeddings, two Transformer layers, and query pooling
15% of valid residues, then 15% of the valid modes within each
Normalized Laplacian coefficients over all valid modes (Smooth-L1)
DNA / RNA
Base, molecule-type, CCD, and validity features; nucleotide-specific two-layer MLP
20% of valid nucleotides
Base and nucleotide-CCD identities (cross entropy)
Ligand
Element, atom-name, molecule-type, CCD, and validity features; ligand-specific two-layer MLP
20% of valid atoms and components
Element, atom-name, and component-CCD identities (cross entropy)
Relations
Ordered structural pair features, molecule-type pairs, and ligand-bond inputs
Sampled valid pairs; supervised bond inputs hidden
Pair distance (Smooth-L1) and bond type (cross entropy)
Appendix
Table 4: Canonical inputs, encoders, masking policies, and reconstruction targets of the multimodal pretraining model. Task-specific overrides are described below.
Figure 7: Ligand-pocket adaptation pipeline. Adaptation of the pretrained residue representation to ligand-binding pocket classification. The native AtomSurf molecular surface and residue graph enter unchanged; in parallel, the frozen pretrained encoder produces residue-level single representations si . The two are concatenated per residue and consumed by the existing AtomSurf graph-input block, so the fused features then follow the original surface–graph encoder and ligand-classification head without further modification. The same downstream path is used for the full and + Ours (w/o TetSphere) variants; the two differ only in the structure-specific input supplied to the frozen pretrained encoder. The right-hand column shows four of the seven cofactor classes.
Figure 8: Test pockets our representation classifies correctly and native AtomSurf does not. TetSphere surfaces are colored by distance to the native ligand (red close, blue far); labels give the true class and the class AtomSurf predicts instead.
Figure 9: PPI residual adaptation pipeline. Residual adaptation of the pretrained residue representation for protein–protein interface prediction. AtomSurf consumes the molecular surface and residue graph of the bound complex and produces the baseline Pinder-Site and Pinder-Pair logits ℓi and ℓLR . In parallel, the frozen pretraining model yields residue-level single representations, from which a site adapter predicts a per-residue correction Δi and a pair adapter predicts a residue-pair correction ΔLR from a symmetric combination of the two projected endpoints. Each correction is added to the corresponding AtomSurf logit, so the AtomSurf branch is left unchanged. The complex shown is barnase–barstar (PDB 1BRS); on the molecular surface, residues within 5 Å of the partner chain are highlighted.
Figure 10: Further predicted interface regions, drawn exactly as in Figure 5 . These three complexes have longer chains than the three shown in the main text ( 216 – 535 , 280 – 281 and 90 – 438 residues, against 97 – 215 ), so more unrelated structure falls in frame behind each patch; the ribbon context is drawn faint and segments that would pass in front of the envelopes are removed. Each panel draws 30 predicted contacts and all 30 are true contacts.
Figure 11: Representation alignment for de novo protein binder design. During training, the fixed multimodal pretraining model provides single and pair representation targets for known target–binder complexes. Learned projectors map the corresponding BoltzGen trunk representations into these target spaces, where the alignment loss is applied. The pretraining model and alignment projectors are used only during training; generation follows the native BoltzGen inference pipeline.
Figure 12: Binders designed by our model on three BoltzGen Challenge Set targets. The designed binder is drawn as ball-and-stick inside its own TetSphere volumes (blue), and the target residues it contacts are shown as TetSphere volumes (red); the remainder of the target is a secondary-structure cartoon.