Learning from Hetero Density for Cryo-EM Protein Reconstruction
Authors: Xu Han, Chaozhuo Li, Xiaowei Yuan, Yuancheng Sun, Kang Liu, Qiwei Ye
Organizations: University of Chinese Academy of Sciences, Beijing, China · Key Laboratory of Complex Systems Cognition and Decision, Institute of Automation, Chinese Academy of Sciences, Beijing, China · Beijing Academy of Artificial Intelligence, Beijing, China · Beijing University of Posts and Telecommunications, Beijing, China · Ant Group, China
Reconstructing protein structures from cryo-electron microscopy (cryo-EM) maps is essential for understanding macromolecular assemblies. Although learning-based methods have improved protein reconstruction, information from hetero components remains underused. Our analysis finds both false predictions and reference protein sites near hetero components; filtering nearby candidates can improve or impair chain construction. We introduce CryoCue, a framework that uses hetero information to guide protein reconstruction. An anchor-supervised detector learns hetero representations across five component classes. Multiscale hetero features guide backbone localization, while predicted hetero candidates condition structure refinement through their class, confidence, and frame-relative geometry. Experiments show that CryoCue improves backbone localization near hetero components and achieves more accurate protein structure reconstruction.
Figures & tables
Figure 1: Hetero context in cryo-EM maps. A, Density assigned to protein and hetero using reference coordinates. B, Local contacts between protein and hetero components in reference structures; gray meshes show experimental density. C, Hetero prevalence and voxel counts across the sample dataset. M denotes million voxels.
Figure 2: Analysis of protein candidates near hetero components. a, Patterns of false protein candidates and protein support near hetero components. b, Effects of candidate filtering on chain construction, shown as improved, degraded, or unchanged.
Figure 3: Overview of CryoCue . a, Hetero-guided protein reconstruction. b, Anchor-supervised detection produces multiscale features and class predictions. c, Coarse attention and fine-scale gating condition backbone localization. d, Predicted hetero candidates condition refinement through their class, confidence, and geometry relative to evolving residue frames.
Structure reconstruction
Density support
Connectivity
Method
TM-score ↑
lDDT ↑
Coverage (%) ↑
RMSD (Å) ↓
Q-score ↑
Break rate (%) ↓
EMProt
0.805
0.740
81.07
0.576
0.604
1.20
E3-CryoFold
0.711
0.512
67.05
1.085
0.369
8.72
ModelAngelo
0.728
0.658
73.22
0.592
0.613
2.61
EModelX
0.805
0.687
82.43
0.842
0.447
6.96
CryoCue (ours)
0.846
0.789
85.02
0.571
0.593
0.69
Table 1: Protein reconstruction performance. Results on the test set. Values are map-level means; bold indicates the best mean.
Overall backbone
Hetero neighborhood
Protein–hetero contacts
Method
CA precision ↑
CA recall ↑
CA precision ↑
CA recall ↑
Precision ↑
Recall ↑
EMProt
96.534
82.675
93.035
86.181
85.646
73.993
E3-CryoFold
86.013
79.818
71.638
76.606
67.762
59.699
ModelAngelo
93.688
81.225
97.511
84.937
87.765
73.145
EModelX
92.946
83.557
75.846
86.832
71.419
64.636
CryoCue (ours)
96.897
85.693
96.271
89.883
86.396
77.877
Table 2: Backbone localization and protein–hetero contact. All scores are percentages; higher is better. Bold indicates the best mean among the compared methods.
Structure reconstruction
Density
Connectivity
Method
TM-score ↑
lDDT ↑
Coverage (%) ↑
RMSD (Å) ↓
Q-score ↑
Break rate (%) ↓
Protein-only
0.780
0.711
77.44
0.609
0.602
1.54
Backbone fusion only
0.819
0.760
82.53
0.588
0.622
1.06
Full w/o anchor supervision
0.747
0.692
75.80
0.715
0.545
1.78
Full w/ fixed hetero geometry
0.840
0.782
84.54
0.573
0.596
0.75
Full model
0.846
0.789
85.02
0.571
0.593
0.69
Table 3: Ablation of hetero representation and conditioning. Variants isolate hetero fusion during backbone localization, anchor-based supervision, and dynamic frame-relative conditioning.
Figure 4: Hyperparameter sensitivity and inference efficiency. Left: TM-score under different candidate caps K and neighborhood radii r , with the other hyperparameter fixed. Right: end-to-end inference time across methods.
Figure 5: Case study of global reconstruction and recovery near a hetero component. a, Global reconstruction of 8IH5 (EMD-35440) by EMProt, ModelAngelo, and CryoCue; the highlighted region is enlarged below. b, Reconstruction near the corresponding hetero component. Magenta markers indicate unrecovered reference protein sites.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Step
Criteria
Database search
Single-particle cryo-EM maps in EMDB and electron-microscopy structures in RCSB, with reported resolution ≤4 Å.
Pair eligibility
A released primary map, an official fitted EMD–PDB association, available map and coordinate files, and at least one protein chain.
Structure validation
Parseable required mmCIF fields and reliable component classification using the Chemical Component Dictionary (CCD) ( Westbrook et al., 2015 ) and entity/polymer metadata. Pairs with unresolved classification conflicts are excluded.
Raw-map quality
Complete downloads; MRC dimensions and voxel spacing consistent with official metadata.
Appendix
Table 4: Data collection and quality control. These checks determine pair eligibility; local label validity is handled separately in Appendix B.4 .
Split or test subset
Pairs
Grid crops
Training
17,820
185,373,427
Validation
300
3,543,489
Test
300
2,029,806
Appendix
Table 5: Dataset split sizes and complete stride-16 crop-grid counts. Grid counts include empty-density regions and boundary crops; they are not counts of independently sampled training examples.
Step
Operation
Coordinate alignment
Use MRC axis order, header origin, and grid starts to express voxels and fitted coordinates in the same physical frame.
Resampling
Resample to isotropic 1 Å spacing using cubic convolution ( a=−0.5 ), preserving map–coordinate alignment.
Normalization
Apply full-map quantile clipping and scaling, as defined below.
Input crops
Extract 483 crops at stride 16. Zero-pad positions outside the map and exclude them from supervision.
Appendix
Table 6: Map preprocessing for the detector and backbone network.
Class
Definition
Anchor
Polymeric nucleic acid
Nucleotides assigned to nucleic acid polymers
C4 ′ per nucleotide residue
Glycan
Sugar residues, including NAG, MAN, and FUC
C1 per sugar residue
Nucleotide/cofactor
Free nucleotides and nucleotide-derived cofactors, including NAD/NADH, FAD/FMN, and SAM/SAH
C4 ′
Other ligand
Other reliably classified non-polymer ligands
Representative heavy atom
Lipid/detergent
Classified lipid and detergent components
Representative heavy atom
Appendix
Table 7: Hetero class definitions and supervision anchors. Named anchors must be uniquely present in the corresponding residue or component.
Split or test subset
Polymeric nucleic acid
Glycan
Nucleotide/cofactor
Other ligand
Lipid/detergent
Training
4,040
3,983
3,859
8,161
1,743
Validation
104
31
68
99
17
Test
110
32
53
98
21
Appendix
Table 8: Pair-level presence of the five hetero classes. A pair is counted once for each class present in its structural labels, so counts across classes are not additive.
Support class
Voxels
Share
Background
474,145,598,844
97.8487%
Protein backbone
3,256,567,203
0.6721%
Protein side chain
5,496,114,368
1.1342%
Polymeric nucleic acid
1,472,613,528
0.3039%
Glycan
50,586,134
0.01044%
Nucleotide/cofactor
9,619,500
0.00199%
Appendix
Table 9: Class composition of the masked structural-support census in the training split, before anchor-target restriction.
Quantity
Value
Analyzed map–structure pairs
1,868
Pairs with nonzero hetero-labeled voxels
1,458 (78.05%)
Protein-associated voxels / full grid
0.709909%
Hetero-associated voxels / full grid
0.116131%
Background voxels / full grid
99.173960%
Hetero / protein-plus-hetero voxels
14.06%
Appendix
Table 10: Full-map statistics for the 1,868-pair preliminary-analysis subset. Voxel fractions are pooled over complete map grids. The final three rows summarize residue-level proximity for 176,963 non-water, non-protein residues retained by the historical parser and are not five-class prevalence estimates.
Class
Pairs
Residues
FP only
FP + protein support
Protein support only
Neither
FP occurrences
TP occurrences
Nucleic acid
264
105,890
78.4%
1.5%
0.1%
19.9%
501,865
1,326
Glycan
396
14,330
72.9%
6.6%
0.6%
20.0%
68,320
910
Nucleotide/cofactor
273
1,172
21.2%
71.2%
6.9%
0.6%
6,881
1,143
Other ligand
751
7,114
50.2%
44.0%
2.6%
3.2%
42,118
4,021
Lipid/detergent
207
6,007
77.3%
20.0%
0.4%
2.3%
40,676
1,246
Appendix
Table 11: Protein candidate patterns near hetero components in the 1,362 audited cases. Percentages are pooled over residue neighborhoods within each analysis group. Protein support includes either a TP candidate or a reference protein site. FP and TP occurrences allow the same candidate to appear in overlapping neighborhoods. Rounding may cause small departures from 100%.
Class
Paired n
Win/loss
Net Δ TP
Removed states/map
Nucleic acid
102
84/17
+4,620
961
Glycan
172
96/70
+687
104
Nucleotide/cofactor
84
40/41
+200
34
Other ligand
357
185/158
+389
45
Lipid/detergent
111
65/44
+523
78
Appendix
Table 12: Effects of candidate filtering on chain construction for the hetero groups defined above. Win/loss reports the numbers of paired cases with increased or decreased C α F 1 at 3 Å; paired n includes unchanged cases. Net Δ TP is the total change in matched protein sites, and removed states/map is the mean number of filtered candidate states in each paired subset.
Stage
Feature shape
Operation
Input
1×483
Normalized density
Scale 1
32×483
Stem and local blocks
Scale 2
64×243
Downsampling and local blocks
Scale 4
128×123
Downsampling and local blocks
Scale 8
256×63
Bottleneck
Global context
384×63
10 Transformer blocks
Appendix
Table 13: Architecture of the hetero detector. Spatial scales are downsampling factors relative to the input crop.
Scale s
Protein Fsp
Hetero Fsh
Fusion
8
256×63
256×63
Global cross-attention
4
128×123
128×123
Window cross-attention
2
64×243
64×243
Gated convolution
1
32×483
64×483
Gated convolution
Appendix
Table 14: Multiscale hetero fusion in the backbone predictor. Shapes give channels and spatial dimensions for a 483 input crop.
Setting
Hetero detector
Backbone predictor
Amino acid predictor
GPUs
8
32
32
Per-device batch
32
16
16
Global batch size
256
512
512
Optimizer updates
225,170
95,231
95,231
Learning rate
10−4
10−4
10−4
Warmup updates
1,000
500
500
Appendix
Table 15: Optimization settings for the voxel-stage models. Learning rates are peak values and training duration is measured in optimizer updates.
Class
Precision (%)
Recall (%)
F 1 (%)
Nucleic acid
95.71
63.08
76.04
Glycan
45.33
15.20
22.77
Nucleotide/cofactor
80.76
16.10
26.85
Other ligand
51.52
46.90
49.10
Lipid/detergent
15.15
13.13
14.07
Appendix
Table 16: Class-wise hetero detection performance. Precision, recall, and F 1 are reported for the five hetero classes.